boxiaolanya2008/tokensqueezer ↗★ 0
@local/tokensqueezer
压缩模型输出和思考块以减少Token消耗 适合需要优化Token使用成本、折叠冗长回答的对话任务。
安装
$
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:boxiaolanya2008/tokensqueezer说明文档
阅读完整 README ↗Configuration
Every key has a default; nothing needs configuring. The floors default to 0, so an ordinary
payload always enters the compressor — and the net-shorter rule above decides whether the result is
kept.
| Key | Default | Meaning |
|---|---|---|
compressOutput | true | Master switch for output compression. |
compressStreamOutput | true | The llm/stream leg — the only one that runs before display. |
compressDurableOutput | true | The session/event backstop. |
compressReasoning | true | Fold reasoning blocks. false skips them before any compressor runs. |
compressToolResults | true | Compress tool results, behind the code guarantee. |
outputMinChars | 0 | Floor for a visible text block. |
reasoningMinChars | 0 | Floor for a reasoning block. |
toolMinChars | 0 | Floor for a tool result. |
minSavedChars | 0 | Worth-it floor. Safe at 0 because of the net-shorter rule. |
segmentMinChars | 200 | Floor for one prose segment when an answer is split around code. |
reasoningElideMinChars | 0 | Compressor-internal floor for reasoning. |
reasoningFoldMinChars | 0 | Compressor-internal fold floor for reasoning. |
answerElideMinChars | 0 | Compressor-internal floor for answers. |
answerFoldFloor | 0 | Compressor-internal fold floor for answers. |
outputTokenCap | unset | Hard ceiling on generated tokens. Too low truncates the answer (finish: max-tokens). Never raises an existing stricter limit. |
reasoningEffort | unset | Reasoning tier to force. Quality tradeoff; applied only when the model advertises the value. |
brevityInstruction | true | Ask for lean answers. The one writing-style change the plugin makes. |
reasoningBrevityInstruction | true | Ask for lean reasoning. |
level | 1 | 0 lossless only, 1 balanced, 2 aggressive. |
relativizePaths | true | Relativise repeated absolute paths. Never on code or diffs. |
internSymbols | false | Intern long repeated identifiers. Never on code or diffs. |
protectCodeFences | true | Keep fenced code out of the semantic layer. |
toolSkip | [] | Tool names whose results are never compressed. |
archMaxChars / archMaxDepth | 1200 / 3 | Bounds for the architecture digest. |
outputTokenCap and reasoningEffort apply to every agent, subagents included — deliberately, so
a delegation cannot quietly outspend the budget it was given.