Meaple-SFKY/dsh-model-orchestrator ↗★ 0
dsh-model-orchestrator
自动发现模型池并智能路由任务的协调器 适合拥有多模型、多提供商,需要根据任务类型自动分发至最佳模型的场景。
安装
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:Meaple-SFKY/dsh-model-orchestrator说明文档
阅读完整 README ↗dsh-model-orchestrator
English | 中文
A generic Model Orchestrator for the DeepSeek Harness. It discovers the models the running harness actually has, works out what each one is evidenced to be good at, and routes each unit of work to the best available one — so you never have to decide which task goes to which model.
It names no model, provider, or vendor anywhere in its selection logic, and it is not specific to any business domain.
What it does
| Concern | How it is handled |
|---|---|
| Model discovery | Reads the live LLM registry at activation, when the adapter topology changes, at most five minutes after the last read, and on demand — the panel's Refresh pool, orchestrate_models { refresh: true }, or GET /state?force=1. A provider's model listing can be a network round trip, so the periodic re-read is lazy: it happens when something next needs the pool, and a failure keeps the pool that already exists. The pool is never hardcoded and never persisted. |
| Capability profiling | Builds a profile per model from authoritative host facts (input modalities, context window, exposed reasoning efforts) plus the provider's own declared description. Nothing is invented. |
| Task matching | Turns a task into a requirement set, then scores every live model against it deterministically. A hard requirement that cannot be evidenced rejects a model instead of silently downgrading. |
| Open capability system | A capability is a generic descriptor, not a domain→model table. Unrecognized domains mint a new descriptor from the task's own vocabulary, which is persisted so the taxonomy genuinely grows. |
| Two modes | Auto infers what the task needs. Guided seeds matching with the capability areas you select for the session. |
| Orchestration | Simple work runs directly. Focused work goes to one specialist subagent. Complex multi-domain work is orchestrated across several, with every expert result returned to the calling agent. |
| Captain | The captain is a role, not a model binding: the task owner that understands, decomposes, dispatches, aggregates, verifies, and closes. Its route is chosen by the matcher from the live pool. |
| UI | A Model Orchestrator settings page, and an Orchestrator board beside Chat and Trajectory showing the session's delegations. Automation by default; everything is adjustable. |
| Persistence | Preferences, learned capability descriptors, and per-route calibrations. Never the model pool. |
| Compatibility | Refuses to activate against an unsupported host, with a precise reason. No silent degradation. |
It is a scheduler, not a task manager
The orchestrator is invoked inside normal task execution. It does not replace, mirror, or compete with DSH's native task handling:
| Concern | Owner |
|---|---|
| Task list, plan, steps, step status | DSH — the orchestrator registers no task tool |
| Progress and progress display | DSH — the orchestrator renders no progress surface and emits no event |
| Session transcript | DSH — the orchestrator appends no custom session event |
| Subagent visibility | DSH — delegated children are ordinary DSH subagents in the normal views |
| Record of what was done | DSH session log — the orchestrator keeps no run history |
| Which model runs a unit of work | Orchestrator |
| Whether a task needs one specialist or several | Orchestrator |
So: todo_write, plan mode, step tracking, and progress rendering all keep working
exactly as they do without this plugin. An expert's answer comes back as the tool result
of the orchestrate_* call, and the agent writes the final answer as usual.
The only live bookkeeping the plugin keeps is a set of abort controllers for delegations currently in flight, so a plugin reload aborts children instead of orphaning them. It is not queryable and is never surfaced as task state.
Why it routes through subagents
The harness exposes exactly one supported seam for choosing a model: agentOptions on a
child agent (ctx.subagents.start). An agent's own route is fixed when it is created,
and there is no per-turn retargeting hook for a running session.
So the orchestrator works with that seam rather than against it:
- Your session's agent stays the user-facing surface and the captain.
- The plugin picks a route per unit of work and spawns a child on it.
- The expert's result returns to the captain as the tool result.
- The captain writes the final answer — the orchestrator never talks to the user.
Install
dsh plugin --profile add github:Meaple-SFKY/dsh-model-orchestrator
Or from a local checkout, as a tarball:
npm pack --pack-destination /tmp
dsh plugin --profile web add /tmp/dsh-model-orchestrator-0.3.0.tgz
Installing the directory itself (add /path/to/dsh-model-orchestrator) does not work, and
fails in a way that looks like a plugin bug rather than an install problem:
pnpm records a link: dependency, so the plugin's real path stays outside the profile and
Node's parent-walk never reaches the profile's node_modules — where the host packages it
imports (@deepseek-ai/dsh-tools, @deepseek-ai/dsh-llm, …) actually live. The load then
fails with Cannot find package '@deepseek-ai/dsh-tools'. A packed tarball is materialized
inside the profile, so resolution works. It also means the profile holds a snapshot: after
editing the checkout, pack and add again.
Once the package is on npm this shortens to the same thing:
dsh plugin --profile add dsh-model-orchestrator
For the record, and so nobody plans around it: the package is not on npm yet, because npm now
requires either 2FA or a bypass-2FA granular token to publish and this account has 2FA disabled —
TOTP enrolment is no longer even possible. A GitHub source needs none of that, which is also how
roughly half of the plugins in the community registry are installed. See
docs/marketplace-submission.md.
The bundle patch mounts one row into the profile's host composition, registers the
orchestrate_* tools into the shared tool registry, contributes one routing-policy
section to the system prompt, and serves the control-panel routes. Restart the profile
after installing so the host picks up the new bundle.
Check the host range first. This release declares engines.dsh = "0.1.5-rc.1" and refuses
to activate against anything else, with the reason and a recovery line. Verify before installing,
or in CI, with:
node scripts/check-compat.mjs
In a checkout with no DSH to check against — a fresh clone, or any CI runner — it says so and
exits 0 rather than reporting an incompatibility it never established, and still validates the
parts that need no host: the declared range, the dsh.engines.dsh mirror, the peer
declarations and compatibility.json. Pass --strict when a missing harness should be fatal,
as it would be in a release gate.
It publishes no service, so it needs no isolate realm, and it only consumes host
capabilities (llm, subagents, tools, systemPrompt).
Use
Nothing to configure. Ask for something and the agent routes it:
"Refactor the parser, then run the benchmarks, then write up what changed."
You can also steer it explicitly:
/model-orchestrator— route one task through the orchestrator, whatever the agent would otherwise have decided. See below.- Settings → Model Orchestrator — mode, capability areas, cost preference, parallelism, route allow/deny lists, live pool, routing preview, capability assignments.
The /model-orchestrator command
Automatic routing is the default, but nothing forces the calling model to route rather
than spawn subagents itself — and a live session was observed delegating four heterogeneous
research units through the native subagent tool, leaving every child on the deployment's
single default model. The command is the explicit switch for exactly that case:
/model-orchestrator 分析这个季度的销售数据并写成一份给管理层看的总结报告,包含趋势图表和三条行动建议
/model-orchestrator status
A handler runs without the command line reaching the model (that is the host's command
contract), so the command delivers the task itself, as an ordinary user message, together
with the instruction to call orchestrate_run and to name per-unit routes when the units
differ in kind. The captain then owns the result as usual.
Honest limits: this makes the intent explicit and reliable to deliver, but the agent still
performs the routing — the command does not bypass the captain, and it does not make
orchestration automatic. Automatic remains the default; this is for when you want it
guaranteed. recordInput: false keeps the task from being logged twice, and the command is
only registered when the deployment mounts the commands service.
It answers in your language. The command's replies, its one-line summary and its input
hint all come from the same locale table the panel uses. The language comes from the durable
locale setting when you have set one — and, when you have not, from the panel telling the
host which language it is rendering. That second source is not a nicety: the harness resolves
a language as "explicit setting → browser detection → en" and never writes the browser-derived
value back, so without it the host cannot know the UI language at all in the default case, and
these strings stayed English inside a Chinese UI. The plugin never writes the setting, only
reads it, and an explicit choice always outranks the browser's.
The harness renders a third-party command's description and hint verbatim — it only translates its own built-in commands — so the plugin re-registers the command when the language changes; that re-registration is the only way its palette entry can follow you. The settings page also states the command, its hint and its usage in the panel's own language, so the copy is discoverable without opening the palette.
Tools
| Tool | Purpose |
|---|---|
orchestrate_run | Analyze, match, delegate every unit, and return all results. The main entry point. |
orchestrate_dispatch | Delegate one self-contained unit to one model. Cheaper and more predictable. |
orchestrate_ask | Ask a completed unit of the same run a follow-up question and get its answer, on the route that did the work. |
orchestrate_plan | Show the routing decision without executing it. |
orchestrate_models | The models that actually exist right now, with the evidence behind each profile. |
orchestrate_capabilities | The capability vocabulary, including anything learned. |
orchestrate_configure | Change preferences: mode, cost, parallelism, routes, reasoning levels, the capability assignment table, dependency handoff, and the sibling-question and review bounds. |
orchestrate_status | Current mode, pool, mappings, the assignment table, the researched facts, which cue groups are still built-in, and the health report. |
Every tool that takes a caller's analysis says where each field belongs, and enforces it
rather than trusting it: a value placed at the top level is reported back in
misplacedArguments with the parameter path to use instead (a route belongs in
units[].route or analysis.modelPreference[].route), an unrecognised argument comes back in
unusedArguments, and a top-level modelPreference/unitModelPreference is folded into
analysis and named in foldedIntoAnalysis. A silently ignored route is not possible — but
the report has to be read, which is why the descriptions name it.
Which models are in the pool
Discovery reads the LLM registry, which lists every model every registered adapter advertises. That is not the same as the set a deployment intends you to use: a profile with two providers mounted commonly advertises the same underlying model twice, and a provider may advertise more than the user enabled.
So the pool is narrowed by two layers, in order:
- The deployment's subagent route policy —
subagentModelSelection.current(), the same exact routes the Settings page shows for subagent model selection. Since every model this plugin runs is a subagent route, this is the authoritative answer. - Your own route preferences —
orchestrate_configure'sallowedRoutes/deniedRoutes, applied on top.
Rules that keep this safe:
- The policy constrains the pool only when the service exists, is enabled, and names at least one route. An absent, disabled, or empty policy means the deployment expressed no preference, and discovery stands — filtering to nothing would silently disable routing.
- Routes are matched exactly as
provider/model. Two providers exposing the same model id are different models and are both kept unless a policy excludes one; ids are never deduplicated, because that would discard a legitimate route. - When the policy would leave nothing routable, that is reported as a problem rather than shown as an empty pool.
The panel says which layer narrowed the pool, so a smaller list reads as a decision rather than a fault:
Showing the 7 route(s) this deployment offers for subagents;
4 advertised route(s) are not selectable.
Providers that do not expose reasoning levels
Not every provider offers a level to pick. DSH reports reasoning in three states and the pool shows which one a route is in:
| State | What it means | What the pool shows |
|---|---|---|
adjustable | The provider exposes levels (low, high, …) | A selector, whose options are the levels it reports |
automatic | The model reasons and the provider drives the depth | automatic — no selector, because there is nothing to select |
none | No reasoning is reported | A dash |
A route in the automatic state is used as-is by default: no level is sent, so the provider
does exactly what it would have done anyway. If you know your provider accepts a level it does not
list, you can set one anyway — orchestrate_configure { reasoningEffort: { "": "high" } } —
and it is accepted, sent, and marked manual in the pool. Nothing can verify it, so a level the
provider rejects fails that delegation with the adapter's own error; that is the trade for not
silently ignoring what you asked for. A route that reports no reasoning at all cannot be given a
level, and a route that lists levels keeps the strict rule: an id that is no longer on its list is
a stale entry and is ignored.
A level is never sent to a route that cannot express it, whatever proposed it. The stored
preference, the calling model's own per-unit request, and a capability descriptor's declared level
are all checked against the destination route; a level it does not advertise is dropped — the route
then resolves its own default — and reported as effortUnavailable on that unit's result. It used
to be sent, on the theory that the caller's judgement outranks the plugin's, and the adapter then
rejected the route and the unit produced no answer at all.
The distinction matters in routing, not only in the panel. A capability that requires reasoning
accepts both adjustable and automatic; a requirement that names a level (for example "must
expose high") needs adjustable, since a level that cannot be selected cannot satisfy it. Setting
a level for a route that reports none is refused when you set it, and ignored if it goes stale —
sending an unsupported level would fail the child outright.
Reasoning level per route
The pool's Reasoning column is a selector, not a label. Its options are the levels the
host actually reports for that route (reasoning.efforts), plus default — which means
"let the model resolve its own level", and shows which one that is whenever the host reports
it (reasoning.defaultEffort). The choice is stored per provider/model and a