Meaple-SFKY/dsh-model-orchestrator ↗★ 0

dsh-model-orchestrator

Generic Model Orchestrator for DeepSeek Harness: discovers the live model pool, profiles real model capabilities, and routes each unit of work to the best available model. Auto and Guided modes, specialist subagents, multi-agent orchestration, a dynamic Captain, and a web control panel. Domain-agnostic. 适合拥有多模型、多提供商,需要根据任务类型自动分发至最佳模型的场景。

패키지
dsh-model-orchestrator
호환성
미검증
Harness peer 범위
0.1.5-rc.1 || 0.1.5-rc.2
버전
0.3.0
라이선스
MIT
최근 업데이트
2026. 9. 14.

설치

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:Meaple-SFKY/dsh-model-orchestrator

dsh-model-orchestrator

English | 中文

A generic Model Orchestrator for the DeepSeek Harness. It discovers the models the running harness actually has, works out what each one is evidenced to be good at, and routes each unit of work to the best available one — so you never have to decide which task goes to which model.

It names no model, provider, or vendor anywhere in its selection logic, and it is not specific to any business domain.


What it does

ConcernHow it is handled
Model discoveryReads the live LLM registry at activation, when the adapter topology changes, at most five minutes after the last read, and on demand — the panel's Refresh pool, orchestrate_models { refresh: true }, or GET /state?force=1. A provider's model listing can be a network round trip, so the periodic re-read is lazy: it happens when something next needs the pool, and a failure keeps the pool that already exists. The pool is never hardcoded and never persisted.
Capability profilingBuilds a profile per model from authoritative host facts (input modalities, context window, exposed reasoning efforts) plus the provider's own declared description. Nothing is invented.
Task matchingTurns a task into a requirement set, then scores every live model against it deterministically. A hard requirement that cannot be evidenced rejects a model instead of silently downgrading.
Open capability systemA capability is a generic descriptor, not a domain→model table. Unrecognized domains mint a new descriptor from the task's own vocabulary, which is persisted so the taxonomy genuinely grows.
Two modesAuto infers what the task needs. Guided seeds matching with the capability areas you select for the session.
OrchestrationSimple work runs directly. Focused work goes to one specialist subagent. Complex multi-domain work is orchestrated across several, with every expert result returned to the calling agent.
CaptainThe captain is a role, not a model binding: the task owner that understands, decomposes, dispatches, aggregates, verifies, and closes. Its route is chosen by the matcher from the live pool.
UIA Model Orchestrator settings page, and an Orchestrator board beside Chat and Trajectory showing the session's delegations. Automation by default; everything is adjustable.
PersistencePreferences, learned capability descriptors, and per-route calibrations. Never the model pool.
CompatibilityRefuses to activate against an unsupported host, with a precise reason. No silent degradation.

It is a scheduler, not a task manager

The orchestrator is invoked inside normal task execution. It does not replace, mirror, or compete with DSH's native task handling:

ConcernOwner
Task list, plan, steps, step statusDSH — the orchestrator registers no task tool
Progress and progress displayDSH — the orchestrator renders no progress surface and emits no event
Session transcriptDSH — the orchestrator appends no custom session event
Subagent visibilityDSH — delegated children are ordinary DSH subagents in the normal views
Record of what was doneDSH session log — the orchestrator keeps no run history
Which model runs a unit of workOrchestrator
Whether a task needs one specialist or severalOrchestrator

So: todo_write, plan mode, step tracking, and progress rendering all keep working exactly as they do without this plugin. An expert's answer comes back as the tool result of the orchestrate_* call, and the agent writes the final answer as usual.

The only live bookkeeping the plugin keeps is a set of abort controllers for delegations currently in flight, so a plugin reload aborts children instead of orphaning them. It is not queryable and is never surfaced as task state.

Why it routes through subagents

The harness exposes exactly one supported seam for choosing a model: agentOptions on a child agent (ctx.subagents.start). An agent's own route is fixed when it is created, and there is no per-turn retargeting hook for a running session.

So the orchestrator works with that seam rather than against it:

  1. Your session's agent stays the user-facing surface and the captain.
  2. The plugin picks a route per unit of work and spawns a child on it.
  3. The expert's result returns to the captain as the tool result.
  4. The captain writes the final answer — the orchestrator never talks to the user.

Install

dsh plugin --profile  add github:Meaple-SFKY/dsh-model-orchestrator

Or from a local checkout, as a tarball:

npm pack --pack-destination /tmp
dsh plugin --profile web add /tmp/dsh-model-orchestrator-0.3.0.tgz

Installing the directory itself (add /path/to/dsh-model-orchestrator) does not work, and fails in a way that looks like a plugin bug rather than an install problem: pnpm records a link: dependency, so the plugin's real path stays outside the profile and Node's parent-walk never reaches the profile's node_modules — where the host packages it imports (@deepseek-ai/dsh-tools, @deepseek-ai/dsh-llm, …) actually live. The load then fails with Cannot find package '@deepseek-ai/dsh-tools'. A packed tarball is materialized inside the profile, so resolution works. It also means the profile holds a snapshot: after editing the checkout, pack and add again.

Once the package is on npm this shortens to the same thing:

dsh plugin --profile  add dsh-model-orchestrator

For the record, and so nobody plans around it: the package is not on npm yet, because npm now requires either 2FA or a bypass-2FA granular token to publish and this account has 2FA disabled — TOTP enrolment is no longer even possible. A GitHub source needs none of that, which is also how roughly half of the plugins in the community registry are installed. See docs/marketplace-submission.md.

The bundle patch mounts one row into the profile's host composition, registers the orchestrate_* tools into the shared tool registry, contributes one routing-policy section to the system prompt, and serves the control-panel routes. Restart the profile after installing so the host picks up the new bundle.

Check the host range first. This release declares engines.dsh = "0.1.5-rc.1" and refuses to activate against anything else, with the reason and a recovery line. Verify before installing, or in CI, with:

node scripts/check-compat.mjs

In a checkout with no DSH to check against — a fresh clone, or any CI runner — it says so and exits 0 rather than reporting an incompatibility it never established, and still validates the parts that need no host: the declared range, the dsh.engines.dsh mirror, the peer declarations and compatibility.json. Pass --strict when a missing harness should be fatal, as it would be in a release gate.

It publishes no service, so it needs no isolate realm, and it only consumes host capabilities (llm, subagents, tools, systemPrompt).

Use

Nothing to configure. Ask for something and the agent routes it:

"Refactor the parser, then run the benchmarks, then write up what changed."

You can also steer it explicitly:

  • /model-orchestrator — route one task through the orchestrator, whatever the agent would otherwise have decided. See below.
  • Settings → Model Orchestrator — mode, capability areas, cost preference, parallelism, route allow/deny lists, live pool, routing preview, capability assignments.

The /model-orchestrator command

Automatic routing is the default, but nothing forces the calling model to route rather than spawn subagents itself — and a live session was observed delegating four heterogeneous research units through the native subagent tool, leaving every child on the deployment's single default model. The command is the explicit switch for exactly that case:

/model-orchestrator 分析这个季度的销售数据并写成一份给管理层看的总结报告,包含趋势图表和三条行动建议
/model-orchestrator status

A handler runs without the command line reaching the model (that is the host's command contract), so the command delivers the task itself, as an ordinary user message, together with the instruction to call orchestrate_run and to name per-unit routes when the units differ in kind. The captain then owns the result as usual.

Honest limits: this makes the intent explicit and reliable to deliver, but the agent still performs the routing — the command does not bypass the captain, and it does not make orchestration automatic. Automatic remains the default; this is for when you want it guaranteed. recordInput: false keeps the task from being logged twice, and the command is only registered when the deployment mounts the commands service.

It answers in your language. The command's replies, its one-line summary and its input hint all come from the same locale table the panel uses. The language comes from the durable locale setting when you have set one — and, when you have not, from the panel telling the host which language it is rendering. That second source is not a nicety: the harness resolves a language as "explicit setting → browser detection → en" and never writes the browser-derived value back, so without it the host cannot know the UI language at all in the default case, and these strings stayed English inside a Chinese UI. The plugin never writes the setting, only reads it, and an explicit choice always outranks the browser's.

The harness renders a third-party command's description and hint verbatim — it only translates its own built-in commands — so the plugin re-registers the command when the language changes; that re-registration is the only way its palette entry can follow you. The settings page also states the command, its hint and its usage in the panel's own language, so the copy is discoverable without opening the palette.

Tools

ToolPurpose
orchestrate_runAnalyze, match, delegate every unit, and return all results. The main entry point.
orchestrate_dispatchDelegate one self-contained unit to one model. Cheaper and more predictable.
orchestrate_askAsk a completed unit of the same run a follow-up question and get its answer, on the route that did the work.
orchestrate_planShow the routing decision without executing it.
orchestrate_modelsThe models that actually exist right now, with the evidence behind each profile.
orchestrate_capabilitiesThe capability vocabulary, including anything learned.
orchestrate_configureChange preferences: mode, cost, parallelism, routes, reasoning levels, the capability assignment table, dependency handoff, and the sibling-question and review bounds.
orchestrate_statusCurrent mode, pool, mappings, the assignment table, the researched facts, which cue groups are still built-in, and the health report.

Every tool that takes a caller's analysis says where each field belongs, and enforces it rather than trusting it: a value placed at the top level is reported back in misplacedArguments with the parameter path to use instead (a route belongs in units[].route or analysis.modelPreference[].route), an unrecognised argument comes back in unusedArguments, and a top-level modelPreference/unitModelPreference is folded into analysis and named in foldedIntoAnalysis. A silently ignored route is not possible — but the report has to be read, which is why the descriptions name it.

Which models are in the pool

Discovery reads the LLM registry, which lists every model every registered adapter advertises. That is not the same as the set a deployment intends you to use: a profile with two providers mounted commonly advertises the same underlying model twice, and a provider may advertise more than the user enabled.

So the pool is narrowed by two layers, in order:

  1. The deployment's subagent route policy — subagentModelSelection.current(), the same exact routes the Settings page shows for subagent model selection. Since every model this plugin runs is a subagent route, this is the authoritative answer.
  2. Your own route preferences — orchestrate_configure's allowedRoutes / deniedRoutes, applied on top.

Rules that keep this safe:

  • The policy constrains the pool only when the service exists, is enabled, and names at least one route. An absent, disabled, or empty policy means the deployment expressed no preference, and discovery stands — filtering to nothing would silently disable routing.
  • Routes are matched exactly as provider/model. Two providers exposing the same model id are different models and are both kept unless a policy excludes one; ids are never deduplicated, because that would discard a legitimate route.
  • When the policy would leave nothing routable, that is reported as a problem rather than shown as an empty pool.

The panel says which layer narrowed the pool, so a smaller list reads as a decision rather than a fault:

Showing the 7 route(s) this deployment offers for subagents;
4 advertised route(s) are not selectable.

Providers that do not expose reasoning levels

Not every provider offers a level to pick. DSH reports reasoning in three states and the pool shows which one a route is in:

StateWhat it meansWhat the pool shows
adjustableThe provider exposes levels (low, high, …)A selector, whose options are the levels it reports
automaticThe model reasons and the provider drives the depthautomatic — no selector, because there is nothing to select
noneNo reasoning is reportedA dash

A route in the automatic state is used as-is by default: no level is sent, so the provider does exactly what it would have done anyway. If you know your provider accepts a level it does not list, you can set one anyway — orchestrate_configure { reasoningEffort: { "": "high" } } — and it is accepted, sent, and marked manual in the pool. Nothing can verify it, so a level the provider rejects fails that delegation with the adapter's own error; that is the trade for not silently ignoring what you asked for. A route that reports no reasoning at all cannot be given a level, and a route that lists levels keeps the strict rule: an id that is no longer on its list is a stale entry and is ignored.

A level is never sent to a route that cannot express it, whatever proposed it. The stored preference, the calling model's own per-unit request, and a capability descriptor's declared level are all checked against the destination route; a level it does not advertise is dropped — the route then resolves its own default — and reported as effortUnavailable on that unit's result. It used to be sent, on the theory that the caller's judgement outranks the plugin's, and the adapter then rejected the route and the unit produced no answer at all.

The distinction matters in routing, not only in the panel. A capability that requires reasoning accepts both adjustable and automatic; a requirement that names a level (for example "must expose high") needs adjustable, since a level that cannot be selected cannot satisfy it. Setting a level for a route that reports none is refused when you set it, and ignored if it goes stale — sending an unsupported level would fail the child outright.

Reasoning level per route

The pool's Reasoning column is a selector, not a label. Its options are the levels the host actually reports for that route (reasoning.efforts), plus default — which means "let the model resolve its own level", and shows which one that is whenever the host reports it (reasoning.defaultEffort). The choice is stored per provider/model and a