mrbeandev/dsh-hypercompact ↗★ 0
dsh-hypercompact
提供基于字节预算的零LLM历史上下文压缩 适合需要精确控制上下文大小、减少Prompt缓存失效并能精确召回历史的用户。
安裝
$
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:mrbeandev/dsh-hypercompact說明文件
閱讀完整 README ↗Configuration
All keys are optional. Unknown keys are rejected at mount, so a typo fails loudly instead of silently using a default.
| Key | Default | Meaning |
|---|---|---|
maxRequestBytes | 5000000 | Compact when the estimated next request body reaches this many bytes. |
targetRequestBytes | 1500000 | Compact down to about this. Must be below maxRequestBytes. The gap is the hysteresis: each compaction costs one prompt-cache miss, so compact in big steps. |
housekeepingRatio | 0.5 | Start housekeeping this fraction of the way from target to trigger (default: at 3.25 MB). Housekeeping offloads older images and trims old tool results in place; no checkpoint. 0 = off. |
housekeepingExcerpt | { head: 2000, tail: 1000 } | Characters a trimmed old tool result keeps. |
maxTokens | 0 | Also compact at this many estimated tokens (0 = off). |
contextRatio | 0.85 | Also compact at this fraction of the routed model's contextWindow (0 = off). The lower of the two token triggers wins. |
retainTurns | 2 | Newest complete turns never compacted. |
retainBytes | 400000 | Newest request bytes never compacted. |
allowIntraTurn | true | When one long autonomous turn holds the bytes, compact its older tool traffic, keeping the turn's human message and the newest retainBytes. |
maxCheckpointBytes / minCheckpointBytes | 600000 / 300000 | Bounds on the checkpoint. Its budget is the space actually free under the target after the retained turns (net of images about to be offloaded), clamped to these bounds. |
pinnedBytes | 30000 | Cap on the pinned block. Trimmed in order: the message index (to the newest 40 lines + a summary), the files list, then the oldest constraints. |
userTextChars | 20000 | Hard cap for ONE pathological human paste: above it, head + tail are kept with a recall pointer. Budget passes never cut human text. |
assistantTextChars | 1500 | Assistant text length in the checkpoint (head + tail); degraded passes never go below 1000. |
toolResultExcerpt | { head: 200, tail: 100 } | Excerpt of each compacted successful tool result. |
toolErrorExcerpt | { head: 600, tail: 400 } | Excerpt of each failed tool result (never dropped by budget passes). |
groupTools | [read, read_image, grep, glob, ls, web_search, web_fetch, recall, recall_search] | Consecutive successful calls to one of these tools collapse into one line. [] = off. |
keyArgChars | 160 | Max characters of one argument shown on a tool-call line. |
largeArgBytes | 2000 | Argument values larger than this are shown as ``. |
keyArgTools | null | null = show key arguments for every tool; or a list of tool names. Other tools show only their argument size. |
keepRecentImages | 2 | Newest inline images kept; older ones become captioned labels when the request is above target. |
maxKeptImageBytes | 2000000 | Byte cap (base64) on the kept images, independent of the count. |
recoverOnTimeout | true | After a request timeout/transport failure with a body above the target, compact and retry once. Context-overflow and HTTP 413 failures are always recovered. |
maxRecoveryRetries | 1 | Recovery compactions per agent until it next goes idle. |
dryRun | false | Log what would be compacted; write nothing. |
auto | true | Automatic triggers. With false, only /compact compacts. |
tools | true | Register the recall and recall_search tools. |
commands | true | Register the /recall and /hypercompact commands. |
recallToolName / searchToolName | recall / recall_search | Tool names, in case another plugin already uses them. |
maxRecallChars | 60000 | Page size of one recall answer (longer output continues with offset). |
maxSearchHits | 30 | Max entries returned by one search. |
statsLog / statsDir | true / null | Append one content-free JSON record per compaction to /.jsonl (default ~/.dsh/hypercompact/). |
modelPolicies | [] | Per-route overrides: [{ provider, model, maxRequestBytes?, targetRequestBytes?, maxTokens?, contextRatio?, retainTurns?, retainBytes?, maxCheckpointBytes?, housekeepingRatio? }]. |
harnessEntry | running dsh | Path of the dsh CLI entry, if it cannot be found from process.argv[1]. |
allowUntestedHarness | false | Load on a DSH release outside the supported range, or one that fails the API contract check. |
Why these defaults
- 5 MB trigger. On a slow uplink (100–250 KB/s), 5 MB uploads in about 20–50 s, inside the common 100 s proxy/CDN timeout even at the low end.
- 1.5 MB target. About 6–15 s per request after compaction, with 3.5 MB of growth before the next compaction (each compaction is one cache miss).
retainTurns: 2,retainBytes: 400000. The current and previous turn stay verbatim, so the agent never loses the work in progress.contextRatio: 0.85. A token safety net for a small-window model whose session stays under the byte trigger.keepRecentImages: 2. The screenshots the agent is currently looking at survive, and older ones, the dominant cost in measured sessions, do not.minCheckpointBytes: 300000. Measured: with a smaller floor, sessions whose retained tail is image-heavy got a tiny budget and elided hundreds of entries; at 300 KB every measured session compiled with degradation ≤ 1 and nothing elided, while still landing near the 1.5 MB target.
On a fast connection with a large-window model you can raise both byte values (for example 20 MB → 6 MB); the token trigger still protects the context window.
Per-model policies. None ship by default, because provider and model IDs are
deployment-specific; provider and model must match your route exactly:
modelPolicies:
# a route behind a slow proxy: keep every request under ~30 s at ~100 KB/s
- { provider: my-gateway, model: my-model, maxRequestBytes: 3000000, targetRequestBytes: 1000000 }
# a 128K-window model: fire the token trigger early
- { provider: deepseek, model: deepseek-chat, contextRatio: 0.7 }