QWE13-ART/dsh-tool-folder0

dsh-tool-folder

Fold the DSH tool surface per request + ChainGuard firewall (high-risk block + exfil-chain detection + anti-obfuscation) + BM25/bge-m3 hybrid tools_search. Shrinks schema tokens 80-90% while keeping selection accuracy. v0.2.0 adds a semantic retrieval leg (local Ollama bge-m3, RRF hybrid) and ChainGuard obfuscation detection (concat-rebuild / -enc / certutil / bitsadmin / homoglyphs).

包名
dsh-tool-folder
版本
0.2.3
最近更新
2026年8月30日

安装

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:QWE13-ART/dsh-tool-folder

dsh-tool-folder

Fold the DSH tool surface down to what each request actually needs. Per-agent core + BM25 dynamic loading + heat feedback + CN-intent aliases. Shrinks per-turn schema tokens ~80-90% while keeping (or improving) tool selection accuracy — the RAG-MCP / BoR / SEP-1576 methodology, implemented as a plain cordis plugin for the DeepSeek Harness host.

Mechanism (verified against DSH 0.1.1-rc.1 sources)

SystemPrompt.assemble() → assembly = { sections, contexts, tools, variables }
ctx.waterfall(scope, "system-prompt/assemble", assembly, ctx, () => assembly)
    ← this plugin listens here and returns {...assembly, tools: filtered}
buildRequest(turn, step, assembly.tools, ...)   ← tools field of the API request

assembly.tools is the ONLY source of the API request's tools. Returning a replaced object is authoritative (model-selection plugin does the same). Execution is NOT gated on the current turn's tools — folded tools still resolve through the full registry — so folding never removes capability.

Feedback and the firewall bind to the execution events, not the prompt assembly:

  • tools/pre-execute — listener receives (exec, next). Return { kind: "deny", reason } to refuse the call, or call next() to allow. exec = { callId, name, arguments, agent, signal } (there is no .deny() method on exec — the contract is the returned decision object).
  • tools/post-execute — listener receives (exec, result, next) and fires once per executed tool; exec.name / exec.arguments feed heat counting, exfil-chain detection, coverage recording and session persistence. (agent/pre-step carries { messages, position, signal } only — no toolCalls — so it cannot drive feedback; verified against dsh-agent-loop.)

Six layers — all implemented

LayerMechanismStatus
L1 schema compressionconservative description/param trimming (no $ref — provider-incompatible), compressEnabled✅ dyn segment only
L2 dynamic loadingper-agent core + BM25 top-K + CN-alias server route✅ core
L3 selection metricsselectionCoverage per session + offline retrievalMetrics✅ recorded to feedback.json
L4 execution sidefolded tools stay executable + heat promotion + tools_search tool
L5 cache disciplinedeterministic ordering (core by config, hot/dyn by name) = byte-stable prefix

Three-segment layout

SegmentContentCost
coreper-agent core + heat-promoted tools, full schemasmall, constant
dynBM25 top-K for the current query + CN-alias server top-Ksmall, per request
foldedeverything else — dropped from the schemazero

An alias hit no longer pulls the whole server in: the matched server's tools are ranked by BM25 (with the matched alias keywords bridged into the query to cross the CN/EN gap) and the top min(3, serverSize) are loaded. When the query yields no lexical signal at all, a stable name-order subset is loaded so the alias still helps.

Optionally (catalogEnabled: true, default) folded tools are listed as a one-line catalog section so the model still knows they exist and can call them by name (execution-side fallback still works).

Install

Official (needs the dsh CLI / pnpm):

dsh plugin --profile desktop add file:E:/DSH-Data/dsh-tool-folder
# or, after publishing: dsh plugin --profile desktop add github:/dsh-tool-folder

Manual fallback: put the package where the profile loader can resolve dsh-tool-folder, then add to profiles//cordis.patch.yml:

- insert:
    - id: tool-folder
      name: 'dsh-tool-folder'
      config:
        enabled: true
        core: [tool-bash, tool-pwsh]
        topK: 6
        hotThreshold: 3

(Manual install is not yet end-to-end verified — the loader resolution path for a locally-added bundle is the one open question.)

Rollback

Set enabled: false (or remove the insert row) and restart DSH. Folded tools were never unregistered, so nothing else changes.

Config

keydefaultmeaning
enabledtruemaster switch
perAgent{}{agentId: {core: [...]}} overrides
core[]always-loaded tools (fallback for unknown agents)
deny[]never inject + refuse execution: exact name, or prefix* matches a whole server. Deny wins over core, include and the catalog too.
topK6BM25 dynamic segment size
hotThreshold3folded tool called N times → auto-promote to core
catalogEnabledtrueappend one-line catalog section for folded tools (fold safety net)
schemaToolEnabledtrueregister the tools_schema meta tool (full parameter schema of one tool by name)
compressLeveloffL1 compression tier: off | light | standard | aggressive. standard trims descriptions to 200 chars + param descriptions to 120; aggressive additionally drops optional parameters (required ⊆ properties guaranteed constructively). Applies to dyn segment only — core/hot never trimmed.
compressEnabledfalsedeprecated compat: compressLevel: off + compressEnabled: true is treated as standard
normalizeDescriptionsfalsedyn-segment description hygiene: injection-marker removal + whitespace normalization + first-sentence keep + 300-char cap
category{}intent routing { "记忆/回忆/remember": ["mcp__viking"], ... }. Key words are /-separated (OR). Category wins over aliases; each matched server loads per-server top-3.
include[]whitelist: exact name or prefix*. Non-empty → only matching tools are injected (meta tools exempt; deny still wins). Execution is NOT gated by include.
toonifyResultsfalsecompact long JSON results after execution (drop empty fields + truncate strings; only when the text block is >2000 chars and parses as JSON)
maxFoldMs50hard cap; beyond this keep the full list
semanticEnabledtruev0.1.8: tools_search semantic leg — BM25 + local bge-m3 (Ollama) RRF hybrid. Chinese intents hit English-only tools (the documented BM25 gap). Index lazily built + disk-cached (~/.dsh/state/semantic-cache.json); Ollama offline/timeout degrades to BM25-only, never slower or worse.
ollamaBasehttp://127.0.0.1:11434Ollama endpoint for the semantic leg
embedModelbge-m3embedding model for the semantic leg
feedbackFile''feedback JSON path (default $DSH_HOME/logs/tool-folder/)
aliasesbuilt-inCN-intent keyword → server prefix table (per-server top-K)

Files

  • lib/bm25.js — zero-dep BM25 (k1=1.2, b=0.75; CN bigrams + EN tokens)
  • lib/semantic.js — v0.2.0 semantic leg: local bge-m3 embeddings (Ollama) + RRF fusion, disk-cached
  • lib/chainguard.js — ChainGuard firewall: literal patterns + exfil-chain detection + deobfuscation
  • lib/obfuscation.js — v0.2.0 anti-obfuscation: concat rebuild, -enc/certutil/bitsadmin shapes, homoglyphs
  • lib/schema.js — L1 tiered compression (compressLevel) + description normalization + lossless-JSON sanitizer
  • lib/category.js — P1-1 intent-category routing (pure function, unit-testable)
  • lib/metrics.js — heat/feedback persistence (fold promotion signals)
  • lib/toonify.js — P2-2 long JSON result compaction (pure function)
  • lib/index.js — assemble hook, three-segment fold, heat feedback, safety
  • cordis.patch.yml — bundle insert declaration
  • test.js — simulated-cordis harness: node test.js

Known limits

  • BM25 cannot cross the CN-query / EN-description language gap on its own; the alias table covers common CN intents and now bridges its EN keywords into the per-server BM25 scoring. A future query-rewrite leg (ARK LLM, proven in the GPT Researcher smart retriever) removes this entirely.
  • Alias hits load the server's top-3 relevant tools, not the whole server. A server whose most relevant tool ranks low on a bridged CN query may miss it — the tools_search meta-tool and heat promotion are the recovery paths.
  • catalogEnabled render behavior needs one runtime verification pass (section render is verified; the catalog text itself is standard).
  • toonifyResults mutation of the result object in the host's post-execute waterfall is not yet proven end-to-end (the listener's return is a gate, not the result body) — the pure function + unit tests are in; runtime propagation needs one verification pass. Default off until then.
  • aggressive compression hides optional parameters from the model, so a model that would have used them can't (execution still validates against the registry's original schema — C6). standard is the safe default tier; aggressive is opt-in.

Changelog

v0.2.3 — fold-ux: search fallback catalog + query expansion + stable results + discovery stats (2026-08-30)

  1. Zero-match fallback catalog (fallbackCatalogSize, default 30): when tools_search finds nothing, it returns a name-sorted lightweight catalog (name + 80-char description) instead of an empty list — the model never concludes a capability doesn't exist just because its wording missed. total reports the real catalog size (not the listed slice); the output carries fallback: true. 0 restores the legacy empty result.
  2. CN→EN query expansion (queryExpandEnabled, default on): a 48-term high-frequency intent table appends English synonyms to the search text (search/find, file, image/ocr, database, deploy, schedule…). Only adds terms, never rewrites the original query; exact-name priority still uses the raw query, so precise lookups are byte-identical to v0.2.2.
  3. Stable results (createSearchCache, 60s TTL / LRU 200): the same query returns the same result set across turns — kills per-turn candidate drift.
  4. Discovery stats (recordDiscoveries): every found tool is counted and persisted (30s-throttled, merged into tool-folder-state.json); on restart the top-5 are logged so low-discovery tools can be identified and their descriptions improved (Anthropic's "monitor discovery rate" loop).
  5. 83 tests green (11 new: ux module).
  6. Three-axis audit (Standards / Spec / compatibility): 1 P1 fixed — warm and discovery state now merge-write the shared state file instead of clobbering each other; P2 hardening — tools_search execute got a whole-body fail-open catch, fallback total semantics clarified, comments corrected to process-level scope.

v0.2.2 — tiered disclosure budget + warm LRU + incremental semantic cache (2026-08-30)

  1. Tiered disclosure budget (disclosureBudget, default 0 = off): caps how many tool names + descriptions are disclosed per request. Over budget, tools are demoted tier by tier (T2 name+desc → T3 name-only → T4 hidden) instead of being dropped wholesale; core/hot/meta tools stay tier-1 and never demote.
  2. Warm LRU (maxWarmTools, default 0 = off): recently used tools are kept visible within a bounded LRU, persisted to state so a restart keeps them warm. 0 keeps the exact legacy injection surface.
  3. Incremental semantic cache: the bge-m3 embedding cache now diffs by content fingerprint — only docs whose text changed get re-embedded, and a model change rebuilds once. A failed refresh never clobbers a good cache.
  4. Config persistence: lib/storage.js reads/writes state with a settings → state → memory fallback chain (never touches cordis.patch.yml).
  5. Explicit fail-open: missingMetaTools gates meta-tool availability — any snapshot oddity returns the full tool list rather than a folded one.
  6. 72 tests green (new: budget 8 / warm 8 / semantic-incremental 5 / storage 7 / fail-open 4).

v0.2.1 — heat decay + exact-name priority (2026-08-30)

  1. Hot heat decay (hotWindowDays, default 3): a tool is promoted to always-visible only when its call count stays inside the sliding window — the old code promoted forever, so any tool used 3 times was permanently resident and the fold silently decayed to zero over time. Legacy feedback.json number entries are read as out-of-window → their stale heat expires naturally. Set hotWindowDays: 0 to keep the old behavior.
  2. Exact-name priority in tools_search: exact tool-name match always ranks first, a ≥4-char query that is a name substring is pulled ahead of BM25/semantic hits (deterministic fallback chain). Results stay deduplicated.
  3. 40 tests green (11 new: window decay / legacy compat / priority order).

v0.2.0 — first npm release (2026-08-30)

This is the first version published to npm. It bundles the v0.1.7 firewall scope fix (exec-only hard block) and the v0.1.8 semantic leg + anti-obfuscation work documented below. The 0.1.x line on npm predates all of it; if you installed from npm before, upgrade to 0.2.0.

The v0.1.7 / v0.1.8 entries below use local internal iteration numbers — the npm 0.1.x packages do NOT contain this work. Everything here ships in 0.2.0 only.

v0.1.8 (internal) — semantic leg + anti-obfuscation (2026-08-30)

  1. tools_search semantic leg: BM25 + local bge-m3 (Ollama) RRF hybrid — Chinese intents hit English-only tools that share no surface tokens (the documented BM25 gap, README Known limits). Doc embeddings are cached to ~/.dsh/state/semantic-cache.json keyed by content fingerprint; the ~10s cold start for 300+ tools happens once. Any failure (Ollama down, timeout, cache unreadable) degrades to pure BM25.
  2. ChainGuard anti-obfuscation leg (lib/obfuscation.js): verdict() now runs three legs — (a) literal patterns on the raw text; (b) de-obfuscated text (quoted concatenation "ne"+"t user" / -join rebuild) re-run through the literal patterns; (c) encoded shapes that no literal pattern can see: -enc base64, certutil -decode chains, bitsadmin /transfer, long base64 payload + exec verb, full-width homoglyphs (curl), combining-mark mangling, hex-escaped bytes. All pure regex/string math, zero deps, sync (pre-execute is a sync callback). Covered by 22 unit tests (all pass); harmless concatenation ("Get"+"-ChildItem") still passes.
  3. New config: semanticEnabled / ollamaBase / embedModel.

v0.1.7 (internal) — ChainGuard false-positive fix (2026-08-30)

  1. P0 firewall scoped to exec tools: tools/pre-execute high-risk hard block (verdict()) now applies ONLY to exec-type tools (isExecTool: pwsh/shell/bash/cmd/exec/run/terminal + run_*/exec_*/ssh_*/wsl prefixes). Previously it ran against EVERY tool's arguments, so content-type tools (write/edit/read — whose arguments are file contents or text) were hard-blocked whenever their payload merely CONTAINED dangerous command literals (e.g. writing a test script, a security analysis, or a doc that quotes dangerous command examples). Verified: content write with such literals was blocked before, is allowed after; exec tools still hard-block real dangerous commands. isExecTool classification covered by a 25-case unit check (all pass).
  2. Scope kept: exfil-chain detection (checkChain, post-execute, warn-only) and deny config are unchanged — they never hard-blocked content tools.

v0.1.6 — settings UI schema (2026-08-25)

  1. Config schema exported (schemastery): every toggle now renders as a native form in the DSH settings UI — no more hand-editing YAML. The schema mirrors DEFAULTS one-to-one (18 fields: enabled / core / deny / topK / hotThreshold / catalogEnabled / compressLevel / compressEnabled / schemaToolEnabled / normalizeDescriptions / category / include / toonifyResults / toolSearchEnabled / firewallEnabled / maxFoldMs / feedbackFile / aliases), with per-field Chinese descriptions and range constraints (topK 0-50, hotThreshold ≥1, maxFoldMs 0-1000). Invalid values are rejected with precise messages ("expected off | light | standard | aggressive but got X"). Same pattern