ADVeRTs13/dsh-auto-model-router ↗★ 1

@adverts13/dsh-auto-model-router

混合模型路由器:声明式规则、成本控制与失败回退链。 适合需按成本或质量自动选模型并配置回退的用户。

套件
@adverts13/dsh-auto-model-router
相容性
待驗證
Harness 依賴範圍
0.1.0-rc.6
版本
0.1.1
授權
MIT
最近更新
2026年8月20日

安裝

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:ADVeRTs13/dsh-auto-model-router

Configuration

Add to your profile's cordis.patch.yml:

- insert:
    - id: dsh-auto-model-router
      name: '@deepseek-ai/cordis-plugin-group'
      group: true
      isolate:
        modelRouter: true
      config:
        - id: dsh-auto-model-router-runtime
          name: '@adverts13/dsh-auto-model-router'
          config:
            # ── four levels (the shipped defaults) ─────────────────────
            # reasoningEffort: off = no thinking, medium/high/max = intensity.
            levels:
              L0:                              # everyday chat
                provider: deepseek-official
                model: deepseek-v4-flash
                reasoningEffort: off
              L1:                              # code & tests
                provider: deepseek-official
                model: deepseek-v4-flash
                reasoningEffort: medium
              L2:                              # writing & review
                provider: deepseek-official
                model: deepseek-v4-flash
                reasoningEffort: high
              L3:                              # complex multi-step
                provider: deepseek-official
                model: deepseek-v4-flash
                reasoningEffort: max

            # Keyword rules only take part in fully-auto mode. Empty by default.
            rules: []

            # ── MODE A: heuristic lock (zero latency) ──────────────────
            heuristic:
              windowSize: 3        # count the last N user inputs
              counting: cjk2       # CJK char = 2, others = 1; or 'tokens'
              thresholds:          # width → level (user-adjustable)
                - maxChars: 10
                  level: L0
                - maxChars: 100
                  level: L1        # lock ceiling; >100 asks the user

            # ── MODE C: keyword scoring (fully-auto) ───────────────────
            scoring:
              weights:             # per-hit score per level
                L1: 1
                L2: 2
                L3: 3
              bands:               # score → level (user-adjustable)
                - maxScore: 0
                  level: L0
                - maxScore: 5
                  level: L1
                - maxScore: 15
                  level: L2
                - maxScore: null
                  level: L3

            # ── Tier 2: cost control (user-configurable) ───────────────
            costControl:
              enabled: true
              mode: balanced        # cost-first | quality-first | balanced
              defaultLevel: L1      # level used when no rule matches
              tokenBudgetPerSession: 0   # 0 = no budget limit (default);
                                        # positive value caps per-session spend
              # What happens when the session budget is exhausted:
              #   silent  — switch to the cheapest level without telling the user
              #   notify  — inject a user-visible notice explaining the switch
              #   ask     — pop a dialog; user may keep the current model
              #             (waives the budget for the rest of the session)
              downgradeBehavior: notify

            # ── LLM classifier (fully-auto fallback; enabled by default) ──
            llmClassifier:
              enabled: true
              model:
                provider: deepseek-official
                model: deepseek-v4-flash
              requestTimeoutMs: 10000

            # Fixed-level mode re-asks every N user inputs.
            askEveryInputs: 3

            # ── Tier 4: failure fallback chain (model-level) ───────────
            fallbackChain:
              - provider: deepseek-official
                model: deepseek-v4-flash

            maxRetries: 1

costControl modes

modedefault level (no rule hit)ambiguity drift
cost-firstlowest configured leveldown (cheaper)
quality-firsthighest configured levelup (stronger)
balancedcostControl.defaultLevelstay put

A rule hit always wins over the cost mode: a level the user explicitly matched is never downgraded by the cost layer.

Budget downgrade is visible

tokenBudgetPerSession defaults to 0 (no budget) — the cost layer never forces a downgrade from budget exhaustion, and downgradeBehavior is inert. Set a positive value to cap spend per session:

  • e.g. tokenBudgetPerSession: 300000: downgrade once the session uses 300k tokens.
  • The downgrade behavior is controlled by downgradeBehavior (default notify):
    • notify: inject a user-visible message explaining the budget is exhausted, which model the remaining requests use, and how to raise the cap.
    • ask: pop a dialog asking "switch to the cheaper model or keep the current one?"; choosing "keep" waives the budget for the rest of the session.
    • silent: switch silently (not recommended — the user wonders why answers got worse).

On-load consultation (no client UI)

On the first session after installing or hot-reloading the plugin, DSH:

  1. Injects a status report message (level→model mapping, cost mode, budget, fallback chain). The language follows settings.yaml's locale.preference (zh/en, default en).
  2. Pops a dialog with three questions (asked once per plugin load):
QuestionOptions
Q1 Tier-1 rulesKeep current rules / View current rules (injects the rule list)
Q2 Tier-2 cost modeKeep current / cost-first / balanced / quality-first (applies immediately)
Q3 Other configSkip / Check a setting (type budget, fallbackChain, classifier, …; injects its current value)

The injected lists and values let the user decide whether to adjust; persistent changes still live in cordis.patch.yml (the messages say so). Headless or no-question-channel setups degrade to report-only and never block.