yangyu666/dsh-jev-prune ↗★ 3

dsh-jev-prune

提供基于Jev判断的上下文压缩与精简工具 适合需要智能裁剪历史工具结果以节省上下文Token的用户。

包名
dsh-jev-prune
兼容性
待验证
Harness 依赖范围
^0.1.5-rc.2
版本
0.1.0
许可证
MIT
最近更新
2026年9月24日

安装

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:yangyu666/dsh-jev-prune

Configuration

KeyDefaultDescription
enabledtrueMaster switch
modeljev-latestJudge model
keepModebudgetLayer 1 decision rule. budget: how much to trim is set by the pressure-gap ratio, which results by Jev's ranking (see below). absolute: the legacy fixed-threshold behaviour
keepThreshold0.5Layer 1: in absolute mode, P(keep) ≥ this means no trimming; in budget mode it is a protection ceiling only (results at or above it never enter the candidate pool)
alwaysTrimRatio0.5Layer 1: fixed trim ratio used only under judgeOn: 'always' (that mode has no pressure signal to derive one from). The budget is this fraction of the candidate pool's total character gain. pressure mode computes the ratio from the gap and ignores this key
volumeBudgetThresholdChars8192⚠️ Deprecated (kept only for compatibility): early budget mode anchored the budget on the volume rule; it now uses the pressure-gap ratio instead, so this key no longer takes effect
keepFloorThreshold / minCandidatesForBudget0.2 / 4budget mode small-population fallback: with fewer than 4 judged candidates, only results with P(keep) < 0.2 are eligible (same degraded-mode shape as layer 2)
budgetMinChars0⚠️ Deprecated (kept only for compatibility): same as above, no longer takes effect
resultExcerptChars240Layer 1: per-result excerpt budget copied into the judge's state (see below); 0 restores the blind ok, N chars line
preserveRecent4Layer 1 leaves the most recent N surface nodes alone
headChars / tailChars600 / 200Layer 1: how many head/tail characters a trim keeps
minCharsToPrune400Layer 1: anything shorter is never trimmed
judgeOn / softLimitpressure / 55%Layer 1: when to judge, and the pressure line
compactReceipts / compactOntrue / pressureLayer 2: switch and pressure line (compactSoftLimit, default 70%)
compactModerelativerelative (recommended) or absolute (with compactThreshold)
compactQuantile0.34The trailing fraction taken on each of the two axes; the intersection is used
compactPreserveRecent1Layer 2 leaves the most recent N surface nodes alone; independent from layer 1's wider recent window
minCandidatesForRelative4Minimum population for relative quantiles; below it the mode degrades to an absolute floor (see below) rather than giving up
floorThreshold / minCandidatesForFloor0.2 / 2Absolute floor used in the degraded mode (materially stricter than compactThreshold) and its minimum sample size
neverCompactToolswrite-type toolsLayer 2 never moves these out; comparison is normalized (Edit ≡ edit)
neverPruneToolsWrite / NotebookEditLayer 1 never touches these. Narrower than the row above on purpose: layer 1 only truncates (reversible, the original stays in the session log), so the arguments of diff-style editors (Edit/ApplyPatch…) are fair game; layer 2 removes the pair outright, so it keeps guarding all of them
compactToolsread-only setAllow-list, non-empty by default (DSH_READONLY_TOOLS: read/glob/grep/list/fetch… plus PowerShell read-only cmdlets such as getchilditem/selectstring). Setting it to [] relaxes the gate to the deny-list only — shell calls then become movable too, which is an explicit opt-in into an unsafe mode
evidenceGuard / evidencePatternstrue / built-in listEvidence guard
compactMinChars / receiptMaxRatio2000 / 0.5Layer 2 economical floors
maxCompactionsPerPass3How many compaction transactions one pass may run. Raised to 3 so a large context converges in a single pass instead of being squeezed across many pre-steps; set to 1 for the old behaviour
judgeMaxRetries / judgeRetryBaseMs2 / 300Retry count and backoff base for judge requests (see below); 0 disables retries
dryRunfalseBoth layers only judge and account; nothing is changed
heartbeatFile''Where to persist state (the host swallows plugin logs, so a file is the only external observation channel)

Retrying judge requests

A single network hiccup used to void the entire round of judging — no candidate got a probability and both layers silently did nothing. Failures are now classified:

FailureHandling
Network error / timeoutRetry with exponential backoff (300ms → 600ms, up to 2 retries by default)
429 / 5xxRetry (server temporarily unavailable)
Other 4xx (401 bad key, 400 malformed request)No retry — immediately fatal; retrying only burns quota
Response missing answersNo retry (a retry would most likely return the same broken body)
External signal already abortedNo retry, and no new request is issued

Batches are isolated too: one failed batch no longer discards the remaining ones, and the count shows up as 失败批次 N 个 in the status report. Only when every batch fails is the round treated as failed.

Counter semantics (clarified in the PR #28 review): client.lastRetries is reset at the start of every ask and is what the status report shows, while client.retries is the lifetime total for the client (useful for "has this client ever had to retry?"). client.requests counts HTTP attempts actually issued, including failed ones, so the identity requests === successful asks + retries holds. Using the lifetime counter as if it described the current pass would make the report show 重试 N 次 forever after a single hiccup, with N only ever climbing.

Token-estimate accuracy

estimateTokens is a heuristic (the plugin ships no tokenizer), but its constants are no longer guesses: they were grid-searched against a real BPE tokenizer over 22 samples (English prose, camelCase identifiers, JSON, Windows and Unix paths, git diffs, Chinese, mixed Chinese/English, code blocks, logs, pure punctuation, hex/UUID, table rows, single glyphs, whitespace), scoring on a weighted fit + holdout objective to avoid overfitting.

Mean absolute error drops from 20.5% to 10.7% (holdout 20.5% → 14.4%), and the direction was corrected: the old formula over-estimated pure English by +37% and Unix paths by +44%, and since both layers use this value in a ratio, it was tightening both gates. The new estimate is essentially unbiased (−0.3%).

npm run check asserts accuracy against a holdout set (5 samples that took no part in the fit, with reference lengths measured from the real tokenizer; current MAE 5.5%): a hard MAE ≤ 15% bound plus a directional assertion that English prose must not be over-estimated. This distinction matters (corrected in the PR #28 review): the first version reused the calibration data itself, which made the assertion a tautology — it could only catch "someone hand-edited the constants", never "the constants overfit the fitting set". With a real holdout, pushing wordSlope to 0.9 jumps the MAE to 38.7% and fails immediately. (Also: the holdout reference lengths must be measured with the real tokenizer, not estimated — 4 of the 5 values I first hand-wrote were off by more than 10%, i.e. the assertion would have been built on wrong numbers.)

Pressure-gate failure direction

Both layers fail closed and in the same direction: when the threshold itself cannot be computed (a ratio soft limit with an unresolvable context window), neither layer acts.

The old behaviour was asymmetric — layer 2 skipped when it could not resolve a threshold, while layer 1 simply fell through and proceeded; more subtly, a missing meter left used at 0, so 0 < threshold was always true and judging ran every single round, i.e. the gate did not exist. For a gate whose purpose is to avoid spending Jev calls, "if we cannot tell, do not spend" is the safe direction.

The boundary (corrected during the PR #28 review): failing closed justifies declining to spend, but it must not turn into silently switching the feature off. When the soft limit is an absolute token count (softLimit: 3000), the threshold comes straight from limit.value and the meter is irrelevant — so if the meter is missing or throws, the gate is simply left un-armed for that pass (reported as 压力门本次不设防 at warn level) and judging proceeds. An earlier revision of this PR required a successful measurement unconditionally, which turned "stop wasting money" into "the first layer never runs again" for any host that does not register tokenMeter — strictly worse than the bug it was fixing. smoke_apply.mjs pins both directions.

Note that tokenMeter is a host-provided service; if your host does not expose it, configure softLimit as an absolute token count (or set judgeOn: 'always' / compactOn: 'always') rather than relying on ratio-based pressure gating.

Host compaction threshold vs. softLimit

Layer 1 does not schedule pruneSession itself. The host's compaction-basic bundle calls it when the host reaches its own pressure threshold (thresholdRatio, 0.8 by default) or on context overflow. This plugin's softLimit controls when Jev judging starts and how large the trimming budget is; it does not replace the host threshold.

In pressure mode, a layer-1 trim therefore needs both conditions:

host calls pruneSession
AND
used tokens exceed softLimit (so the pressure-gap budget is greater than zero)

Keep softLimit at or below the host's thresholdRatio unless the delayed behaviour is intentional. For example, with host thresholdRatio: 0.8 and plugin softLimit: 90%, host calls between 80% and 90% produce a zero plugin budget; trimming starts only after usage reaches 90%. With the default softLimit: 55%, judging is ready before the host's normal 80% compaction call.

For @deepseek-ai/dsh-llm-deepseek@0.1.5-rc.2, configure a smaller context window on the matching model entry:

- id: llm-deepseek
  config:
    models:
      - id: deepseek-flash
        contextWindow: 10000

Setting only defaultContextWindow does not override catalog models that already carry their own contextWindow; the model entry wins. If the plugin cannot resolve the effective window, ratio-based gates stop and report the reason instead of guessing.

Layer 1: pressure-quantile trimming

The layer-1 decision used to be a bare fixed threshold: keep = P(keep) ≥ 0.5. Live-host measurement broke that assumption: every judged candidate scored below 0.5 (42/42 in a 132k-token session, 5/5 in a short one; median ≈ 0.13–0.17). Jev's probabilities live in a narrow band — the exact trap layer 2 had already escaped by switching to relative quantiles, except nobody applied the lesson to layer 1. Under the fixed threshold, the first layer's real-world behaviour was "trim everything that was judged", including results the session still needed.

budget mode (the default) decouples the two questions:

  • How much to trim comes from the pressure-gap ratio: ratio = (used − threshold) / window, computed automatically each judge pass (the fraction of the context window that is over the soft limit); the budget is ratio × the pool's total character savings. Zero gap ⇒ trim nothing; the closer to the ceiling, the more is trimmed.
  • Which results comes from Jev: candidates are sorted by P(keep) ascending and trimmed until the budget is met. Results with P(keep) ≥ keepThreshold (0.5) are a protection ceiling and never enter the pool; candidates whose savings are already counted stop there — the rest are recorded as 预算已用尽 (keptByBudget) rather than silently kept or trimmed.
  • With a small population (< minCandidatesForBudget), the mode degrades to an absolute floor (keepFloorThreshold, 0.2) instead of inventing a ranking from 2–3 samples — the same degraded-mode shape layer 2 uses.

The old behaviour stays available as keepMode: 'absolute'.

Result excerpts (giving the judge eyes)

The judge's state used to describe every tool result as ok, 16489 chars (内容省略) — the judge knew that something big existed but not what was in it. Blind judging plus a fixed threshold degenerates into "trim whatever is large".

With resultExcerptChars (default 240), each result line in the state carries a bounded excerpt. Lines are picked by informativeness, not position, because the two naive rules both failed a real-session A/B:

  1. error/evidence-pattern lines (the obvious candidate), plus
  2. salient lines: constant identifiers (THRESHOLD_DISCOUNT_PCT, E2001_BASE_IMAGE), assignments/keys (timeout = 4800), file paths — the critical config line in a 20 KB module is neither at the head nor an error, and rule 1 alone missed it (measured: judging outcomes identical to no excerpt at all), plus
  3. a middle-line fallback for pure-prose results (the middle is exactly what "cut the middle" loses).

The excerpt is hard-bounded per result and counted against the state budget, so it cannot blow up the request size. One trade-off to know: excerpts share the fixed maxStateTokens budget with history lines — at ~70 tokens per excerpted result, a 100-result session spends ~28% of the default 25k-token budget on excerpts, and the squeeze logic compensates by dropping more history lines. If you run very long sessions, raise maxStateTokens (Jev's ceiling is 32k) or lower resultExcerptChars rather than disabling excerpts entirely. Note the interaction with budget mode: excerpts shift probabilities; only the ranking-based decision converts better information into different trimming. Under a fixed threshold both A/B arms behaved identically — the excerpt's value presupposes the ranking rule.

Judgement observability (heartbeat)

The heartbeat now records decision evidence, not just counters — because both the 0.5-threshold failure above and the upstream #25–#29 regressions were invisible in a stats-only heartbeat:

FieldContent
keepLayer-1 decision rule as configured (mode, ceiling, floor, budget parameters)
gateLast pressure-gate evaluation: used / resolved window / threshold / skip + reason / candidate count
probSummaryDistribution of P(keep): p10–p90, mean, counts above/below the ceiling
probSamplesLast 200 raw probabilities (for histograms)
lastJudgePass.rowsPer-candidate detail: seq, tool, chars, prob, effectProb
stats.preStepEvents / stats.judgePassSkipped + lastJudgeSkipReasonDistinguishes "the event never fired" / "no candidates" / "gate skipped" — three failures that used to look identical from outside

One structural fix came out of this: the heartbeat is merge-written, but the pre-step hook itself never called writeHeartbeat — so gate-skipped passes left the file frozen at the boot snapshot (bootedAt == now) and every skipped path was unobservable. The hook now persists after every step.

Out-of-range configuration

Every numeric option has a valid range, and out-of-range values are never passed through to the runtime:

  • Through the Config schema (the host's normal load path) → a ValidationError is thrown. A loud refusal.
  • Without schema normalization (a config object injected directly by cordis.patch.yml, or the PLUGIN_CFG built by the smoke test) → resolveConfig falls the value back to its default (not to the nearest bound, because "how far off was it" isn't interpretable), and records a configWarnings entry in the status report and heartbeat.

These are the real consequences, all of which used to happen silently:

SettingBehavior before the fix
preserveRecent = -5lastAllowed grew instead of shrink