boxiaolanya2008/tokensqueezer ↗★ 0

@local/tokensqueezer

压缩模型输出和思考块以减少Token消耗 适合需要优化Token使用成本、折叠冗长回答的对话任务。

包名
@local/tokensqueezer
兼容性
待验证
版本
0.1.0
最近更新
2026年9月30日

安装

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:boxiaolanya2008/tokensqueezer

Configuration

Every key has a default; nothing needs configuring. The floors default to 0, so an ordinary payload always enters the compressor — and the net-shorter rule above decides whether the result is kept.

KeyDefaultMeaning
compressOutputtrueMaster switch for output compression.
compressStreamOutputtrueThe llm/stream leg — the only one that runs before display.
compressDurableOutputtrueThe session/event backstop.
compressReasoningtrueFold reasoning blocks. false skips them before any compressor runs.
compressToolResultstrueCompress tool results, behind the code guarantee.
outputMinChars0Floor for a visible text block.
reasoningMinChars0Floor for a reasoning block.
toolMinChars0Floor for a tool result.
minSavedChars0Worth-it floor. Safe at 0 because of the net-shorter rule.
segmentMinChars200Floor for one prose segment when an answer is split around code.
reasoningElideMinChars0Compressor-internal floor for reasoning.
reasoningFoldMinChars0Compressor-internal fold floor for reasoning.
answerElideMinChars0Compressor-internal floor for answers.
answerFoldFloor0Compressor-internal fold floor for answers.
outputTokenCapunsetHard ceiling on generated tokens. Too low truncates the answer (finish: max-tokens). Never raises an existing stricter limit.
reasoningEffortunsetReasoning tier to force. Quality tradeoff; applied only when the model advertises the value.
brevityInstructiontrueAsk for lean answers. The one writing-style change the plugin makes.
reasoningBrevityInstructiontrueAsk for lean reasoning.
level10 lossless only, 1 balanced, 2 aggressive.
relativizePathstrueRelativise repeated absolute paths. Never on code or diffs.
internSymbolsfalseIntern long repeated identifiers. Never on code or diffs.
protectCodeFencestrueKeep fenced code out of the semantic layer.
toolSkip[]Tool names whose results are never compressed.
archMaxChars / archMaxDepth1200 / 3Bounds for the architecture digest.

outputTokenCap and reasoningEffort apply to every agent, subagents included — deliberately, so a delegation cannot quietly outspend the budget it was given.