Hubert-hwk/dsh-loop-doctor--packages-loop-doctor-advisor ↗★ 0

@hubert-hwk/dsh-loop-doctor-advisor

Turns loop-doctor detour signals into human-reviewable optimization suggestions (change prompt / add tool / tune skill) with deterministic replay verification. 适合需要审查代理绕路并验证修复建议的开发者,依赖检测器信号。

Package
@hubert-hwk/dsh-loop-doctor-advisor
Compatibility
Unverified
Harness peer range
^0.1.0-rc.8
Cordis peer range
^4.0.1
Version
0.1.0-rc.5
License
MIT
Last updated
Aug 20, 2026

Install

This plugin has no verified bundle, or compatibility checks failed. Read the repository notes first. Read the full README ↗

@deepseek-ai/dsh-loop-doctor-advisor

中文 | English

The advice plane of the loop-doctor family: one DetourSignal → one human-reviewable Suggestion, verified by replaying the session log under the proposed fix.

  • Subscribes to loop-doctor/signal and emits loop-doctor/suggestion (suggestion + measured ReplayOutcome).
  • analyzeAndSuggest(events, sessionId, patterns?, options?, replayCaps?) runs the whole pipeline on demand — the function the diagnose tool calls.
  • replay.ts simulations are counterfactuals over the same log:
    • repeatReminderSimulation(threshold) — calls saved by a repeat-reminder (replay-verified badge).
    • retryBudgetSimulation(maxRetries) — delayMs saved by a retry budget (replay-verified).
    • errorProneToolCallsSimulation(maxErrors) — calls saved by dropping a repeatedly-failing tool (replay-verified).
    • eventCountCapSimulation(cap, …) — events saved by capping compactions / injections / delegations at a fixed policy cap (simulated badge: count arithmetic under a declared deployment policy, never derived from the observed count, so the saving scales with the actual excess).
    • promptBloatSimulation() — chars saved by freezing the prompt at its first observed length (simulated). A fix is replay-verified only when a behavior-modeling simulation shows a measurable improvement on the exact log; policy-cap outcomes carry simulated: true and are never counted as replay-verified.
  • fix-patterns.jsonl is the sediment: the durable "detour → fix" mapping, loaded via loadFixPatternsFile() (plugin config fixPatternsFile) or the built-in BUILTIN_FIX_PATTERNS when absent.
  • Accepted-fix injection (P2): sedimentFile points at the apply plugin's accepted-fixes.jsonl; accepted fixes merge into the pattern set at startup AND live — this plugin provides ctx.loopDoctorPatterns.reload(), which the apply plugin calls after every sediment write, so a validated fix is what gets suggested next time in the same process. Accepted per-context (sub) patterns win over the base ones.

Config

- id: loop-doctor-advisor
  name: '@deepseek-ai/dsh-loop-doctor-advisor'
  config:
    fixPatternsFile: ''   # optional JSONL override; empty ⇒ built-ins
    sedimentFile: ''      # optional accepted-fix JSONL (written by the apply plugin)
    replayCaps:           # FIXED policy caps for simulated context-thrash fixes
      compaction: 2       #   (default; below the detection threshold of 3)
      inject: 10          #   (default; below the detection threshold of 15)
      delegation: 15      #   (default; below the detection threshold of 20)

Safety: confidence is capped at 0.9; patterns are verifiedBy: 'replay' or 'expert' (never presented as proven); nothing is applied automatically.

See the family README for the paper mapping.

Model Experience

Suggestions surfaced only via the diagnose tool

What the model sees

The advisor itself adds nothing to the conversation. When the agent calls loop_doctor_diagnose, the tool's rendered report — built from this package's suggestions — is the only surface, with every suggestion labeled [replay-verified] or [simulated policy counterfactual] so a policy cap cannot be misread as a measured saving.

Suggestion line as rendered by the tool
- suggestion [retry-storm-repeat] (prompt, 80%): Stop repeating identical calls — Add a system-prompt rule (and mount repeat-tool-reminder) [replay-verified]
Token effect

Zero for the plugin itself; the diagnose report's tokens ride the tool result only when the model calls the tool.

KV Cache effect

None from the plugin. The tool result appends after the reusable request prefix and invalidates nothing.

Known Limitations and Deferred Work

  • Replay is a proxy re-derivation, not a rerun. Simulations recompute a metric under a rule change on the same log; they never re-run the loop, so a fix whose benefit depends on changed model behavior is out of scope.
  • Policy caps are declared, not learned. replayCaps are deployment constants; deriving them from the observed count would make the saving tautological (rejected by design).
  • Confidence values are hand-calibrated. The 0.55-0.9 range encodes judgment about rule strength, not a learned probability.
  • No suggestion dedup across sessions. The same detour class in a new session yields a fresh suggestion; sediment acceptance is the only per-context memory.