Hubert-hwk/dsh-loop-doctor--packages-loop-doctor-detector ↗★ 0

@hubert-hwk/dsh-loop-doctor-detector

Reads dsh's SessionEvent log and detects agent detours (retry storms, tool misuse, context thrash, prompt bloat, effect churn) with deterministic rules — no LLM involved. 适合需只读监控代理异常行为并输出信号的开发者,无需模型调用。

패키지
@hubert-hwk/dsh-loop-doctor-detector
호환성
미검증
Harness peer 범위
^0.1.0-rc.8
Cordis peer 범위
^4.0.1
버전
0.1.0-rc.5
라이선스
MIT
최근 업데이트
2026. 8. 20.

설치

검증된 bundle이 없거나 호환성 검사에 실패했습니다. 먼저 저장소 설명을 읽어 주세요. 전체 README 읽기 ↗

@deepseek-ai/dsh-loop-doctor-detector

中文 | English

The observation plane of the loop-doctor family: a read-only listener plugin that turns the session/event firehose into deterministic DetourSignals.

  • Subscribes to session/created / session/event / session/disposed and folds every event into a per-session SessionAnalysis (WeakMap-keyed, die-with-session, no teardown effect needed).
  • Emits loop-doctor/signal at each turn/end (and at adoption / disposal), exactly once per (kind, turn, sub) identity.
  • Hot-reload safe: adopting an already-live session replays its log.
  • analyzeLog(events, sessionId, options) is the same engine in one shot — pure, deterministic, replayable. The rules live in patterns.ts as pure functions with no ctx dependency.

Config

All detector thresholds are validated plugin config (schemastery .default(), re-checked fail-loud in apply):

- id: loop-doctor-detector
  name: '@deepseek-ai/dsh-loop-doctor-detector'
  config:
    retryStormMin: 2        # llm/retry events in one turn
    repeatToolMin: 3        # consecutive identical tool/call run length
    toolErrorMin: 2         # tool/result errors before misuse is eligible
    compactionThreshold: 3  # compaction/start per session
    injectThreshold: 15     # injected (non-user) user/message per session
    delegationThreshold: 20 # subagent/descriptor + hook/invoked per session
    promptBloatSteps: 3     # consecutive growing request/header snapshots

Event → signal mapping

  • llm/retry (per turn) → retry-storm (metric: retry count).
  • tool/call runs of identical canonical arguments → retry-storm, sub: 'repeat:' (metric: run length). Arguments are canonicalized by deep key-sort, so property order never splits a run.
  • tool/result.error whose failure recurs in a later turn → tool-misuse (sub: ), attributed to the last failing turn. A later successful use of the tool is not misuse — only a recurring failure is. The evidence includes only the failed reuse attempts (calls whose results errored after the tool's first error); successful reuses never appear. (Attribution pairs each result with its call via the logged callId, bounded at MAX_CALL_NAMES entries so long sessions stay flat.)
  • compaction/start / injected user/message / delegation markers → context-thrash (sub: compaction | inject | delegation, session-level).
  • Strictly-growing request/header system lengths → prompt-bloat.
  • tool/result marked as a failure — isError: true or the structured error field (a plain throw sets the flag but no error field, so the modeled isError is the reliable signal) — coupled with llm/retry in the same turn → effect-churn.

Every DetourSignal carries its evidence as direct references into the session log — replay session.events to audit a finding.

See the family README for the paper mapping and safety baseline.

Model Experience

Presence of the detector

What the model sees

Nothing. The detector adds no tool, no prompt section, and no message; a session with it mounted is byte-identical on the model surface to one without.

Model surface unchanged
(no diff: the detector emits harness events only)
Token effect

Zero. The detector never touches the prompt or any message content.

KV Cache effect

None. No model-visible bytes change, so no KV-cache entries are invalidated.

Known Limitations and Deferred Work

  • Exact-match detour rules only. Argument canonicalization is a deep key-sort, so near-identical variants (a tweaked path, extra whitespace inside a value) never chain; fuzzy matching is rejected pending evidence of need.
  • Turn-attributed findings are per-session. The fold state dies with the session (WeakMap), so a detour pattern spanning sessions is only visible through session-level counters (compaction/inject/delegation).
  • Prompt-bloat detection is strictly monotone. Alternating growth and shrink never triggers; a smoothed trend needs a defined baseline first.
  • No cross-subagent correlation. A parent and its subagent repeating the same failing call produce independent signals; combining them needs a family-level view that does not exist yet.