ADVeRTs13/dsh-auto-model-router ↗★ 1
@adverts13/dsh-auto-model-router
混合模型路由器:声明式规则、成本控制与失败回退链。 适合需按成本或质量自动选模型并配置回退的用户。
安裝
$
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:ADVeRTs13/dsh-auto-model-router說明文件
閱讀完整 README ↗Configuration
Add to your profile's cordis.patch.yml:
- insert:
- id: dsh-auto-model-router
name: '@deepseek-ai/cordis-plugin-group'
group: true
isolate:
modelRouter: true
config:
- id: dsh-auto-model-router-runtime
name: '@adverts13/dsh-auto-model-router'
config:
# ── four levels (the shipped defaults) ─────────────────────
# reasoningEffort: off = no thinking, medium/high/max = intensity.
levels:
L0: # everyday chat
provider: deepseek-official
model: deepseek-v4-flash
reasoningEffort: off
L1: # code & tests
provider: deepseek-official
model: deepseek-v4-flash
reasoningEffort: medium
L2: # writing & review
provider: deepseek-official
model: deepseek-v4-flash
reasoningEffort: high
L3: # complex multi-step
provider: deepseek-official
model: deepseek-v4-flash
reasoningEffort: max
# Keyword rules only take part in fully-auto mode. Empty by default.
rules: []
# ── MODE A: heuristic lock (zero latency) ──────────────────
heuristic:
windowSize: 3 # count the last N user inputs
counting: cjk2 # CJK char = 2, others = 1; or 'tokens'
thresholds: # width → level (user-adjustable)
- maxChars: 10
level: L0
- maxChars: 100
level: L1 # lock ceiling; >100 asks the user
# ── MODE C: keyword scoring (fully-auto) ───────────────────
scoring:
weights: # per-hit score per level
L1: 1
L2: 2
L3: 3
bands: # score → level (user-adjustable)
- maxScore: 0
level: L0
- maxScore: 5
level: L1
- maxScore: 15
level: L2
- maxScore: null
level: L3
# ── Tier 2: cost control (user-configurable) ───────────────
costControl:
enabled: true
mode: balanced # cost-first | quality-first | balanced
defaultLevel: L1 # level used when no rule matches
tokenBudgetPerSession: 0 # 0 = no budget limit (default);
# positive value caps per-session spend
# What happens when the session budget is exhausted:
# silent — switch to the cheapest level without telling the user
# notify — inject a user-visible notice explaining the switch
# ask — pop a dialog; user may keep the current model
# (waives the budget for the rest of the session)
downgradeBehavior: notify
# ── LLM classifier (fully-auto fallback; enabled by default) ──
llmClassifier:
enabled: true
model:
provider: deepseek-official
model: deepseek-v4-flash
requestTimeoutMs: 10000
# Fixed-level mode re-asks every N user inputs.
askEveryInputs: 3
# ── Tier 4: failure fallback chain (model-level) ───────────
fallbackChain:
- provider: deepseek-official
model: deepseek-v4-flash
maxRetries: 1
costControl modes
| mode | default level (no rule hit) | ambiguity drift |
|---|---|---|
cost-first | lowest configured level | down (cheaper) |
quality-first | highest configured level | up (stronger) |
balanced | costControl.defaultLevel | stay put |
A rule hit always wins over the cost mode: a level the user explicitly matched is never downgraded by the cost layer.
Budget downgrade is visible
tokenBudgetPerSession defaults to 0 (no budget) — the cost layer never
forces a downgrade from budget exhaustion, and downgradeBehavior is inert.
Set a positive value to cap spend per session:
- e.g.
tokenBudgetPerSession: 300000: downgrade once the session uses 300k tokens. - The downgrade behavior is controlled by
downgradeBehavior(defaultnotify):notify: inject a user-visible message explaining the budget is exhausted, which model the remaining requests use, and how to raise the cap.ask: pop a dialog asking "switch to the cheaper model or keep the current one?"; choosing "keep" waives the budget for the rest of the session.silent: switch silently (not recommended — the user wonders why answers got worse).
On-load consultation (no client UI)
On the first session after installing or hot-reloading the plugin, DSH:
- Injects a status report message (level→model mapping, cost mode, budget,
fallback chain). The language follows
settings.yaml'slocale.preference(zh/en, defaulten). - Pops a dialog with three questions (asked once per plugin load):
| Question | Options |
|---|---|
| Q1 Tier-1 rules | Keep current rules / View current rules (injects the rule list) |
| Q2 Tier-2 cost mode | Keep current / cost-first / balanced / quality-first (applies immediately) |
| Q3 Other config | Skip / Check a setting (type budget, fallbackChain, classifier, …; injects its current value) |
The injected lists and values let the user decide whether to adjust;
persistent changes still live in cordis.patch.yml (the messages say so).
Headless or no-question-channel setups degrade to report-only and never block.