Asheblog/dsh-ollama-cloud ↗★ 0

dsh-ollama-cloud

Ollama Cloud provider for DeepSeek Harness: one-click provider setup with per-model reasoning-effort control and account usage in the Models page and sidebar. 适合使用Ollama Cloud模型并需要精细调节推理强度的用户。

パッケージ
dsh-ollama-cloud
互換性
未検証
Harness ピア範囲
>=0.2.0-rc.1
Cordis ピア範囲
>=4.0.4 <5.0.0
バージョン
0.2.0
ライセンス
MIT
最終更新
2026/09/29

インストール

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:Asheblog/dsh-ollama-cloud

ドキュメント

README 全文を読む ↗

Usage

  1. Pick an ollama-cloud model in the composer's model menu (or /model).
  2. Pick an effort in the same menu. The offered levels depend on the model, and the default comes from Ollama's own model metadata.
  3. Thinking streams into the transcript like any other provider. The effort is stored per session; a running step keeps the level it started with.

Shipped models and levels

Snapshot of live metadata (/api/tags + /api/show) taken 2026-09-29:

ModelContextVisionLevelsDefault
deepseek-v4.1-flash1,048,576✔Off / Low / High / MaxHigh
deepseek-v4-pro:08131,048,576—Off / Low / High / MaxLow
kimi-k31,048,576✔Off / Low / High / MaxMax
kimi-k2.6262,144✔Off / HighHigh
kimi-k2.7-code262,144✔Off / HighHigh
glm-5.31,048,576—Low / High / MaxMax
glm-5.3-flash1,048,576✔Low / High / MaxMax
glm-5.21,048,576—Off / High / MaxHigh
gpt-oss:120b131,072—Low / Medium / HighMedium
gpt-oss:20b131,072—Low / Medium / HighMedium
minimax-m3512,000✔Off / Low / Medium / High / Max—
minimax-m2.7196,608—HighHigh
nemotron-3-ultra262,144—Off / HighHigh
nemotron-3-super262,144—Off / HighHigh
nemotron-3-nano:30b262,144—Off / HighHigh
gemma4:31b262,144✔Off / HighOff
mistral-large-3:675b262,144✔(no thinking control)—

Models whose metadata is a boolean switch (kimi-k2.6, gemma4:31b, …) expose Off and High only — on the wire, High means "thinking on". minimax-m2.7 reports [true], so thinking cannot be switched off and only High is offered. minimax-m3 reports the thinking capability without a level ladder, so it takes the standard ladder and claims no default; the same applies to any model neither the catalog nor your configuration describes.

Refreshing the catalog

Ollama retires cloud models (a retired id answers HTTP 410). The shipped catalog is a snapshot, not an authority:

  • Discover from the endpoint: any surface calling llm/discoverModels lists what the endpoint currently serves, with context windows and input modalities. Only models /api/show can fully describe become candidates.
  • Edit by hand: override models in the plugin configuration (below). enabled: false hides a shipped model.

Cloud usage in the UI

The plugin ships a browser half that renders in the host's own surfaces — no separate page, and no provider-UI shell plugin needed:

  • Models page card: inside the Ollama Cloud row (Settings → Models), one meter per billing window with its remaining share and reset line, the primary window's per-model request counts, the write-only API key field, and a refresh button.
  • Sidebar row: a compact remaining-quota line under the session list; click it for the per-window detail. It refreshes when the sidebar mounts and every 15 minutes after that.

The card reads usage through this plugin's own host channel (/ollama-cloud), so the API key stays on the host and never reaches the browser. A local or self-hosted endpoint answers 404 on /usage; that renders as "this endpoint does not report cloud usage" rather than an error, and the last good snapshot keeps showing instead of disappearing.

The usage channel needs the composition's connection row to inject webServer; this bundle's patch adds that (profiles without a connection row skip it). After installing or updating the plugin, restart the harness once — until then the card says so itself.

Cloud usage in the UI

The plugin ships a browser half that renders in the host's own surfaces — no separate page, and no provider-UI shell plugin needed:

  • Models page card: inside the Ollama Cloud row (Settings → Models), one meter per billing window with its remaining share and reset line, the primary window's per-model request counts, the write-only API key field, and a refresh button.
  • Sidebar row: a compact remaining-quota line under the session list; click it for the per-window detail. It refreshes when the sidebar mounts and every 15 minutes after that.

The card reads usage through this plugin's own host channel (/ollama-cloud), so the API key stays on the host and never reaches the browser. A local or self-hosted endpoint answers 404 on /usage; that renders as "this endpoint does not report cloud usage" rather than an error, and the last good snapshot keeps showing instead of disappearing.

The usage channel needs the composition's connection row to inject webServer; this bundle's patch adds that (profiles without a connection row skip it). After installing or updating the plugin, restart the harness once — until then the card says so itself.

Configuration

Every field is editable from the plugin settings page and overridable per row in the profile's cordis.patch.yml:

- id: llm-ollama-cloud
  name: 'dsh-ollama-cloud'
  config:
    apiKeyEnv: OLLAMA_API_KEY        # credential reference; empty string = no auth (local servers)
    baseURL: https://ollama.com/api  # native API root; local Ollama: http://localhost:11434/api
    maxTokens: 32768                 # output cap for models that declare none
    defaultContextWindow: 262144     # context fallback for models that declare none
    streamIdleTimeoutMs: 300000      # maximum idle time between stream reads
    requestTimeoutMs: 15000       # per-attempt budget for the non-chat requests (discovery, web search/fetch)
    retryPolicy:                     # executed by dsh-llm-retry
      mode: normal
      maxRetries: 5
    models:                          # merges over the built-in catalog by id; new ids append
      - id: gpt-oss:20b
        contextWindow: 131072
        reasoningEfforts: { off: none, low: low, medium: medium, high: high }
        defaultEffort: medium
      - id: brand-new-model          # a model newer than the snapshot
        name: Brand New
        contextWindow: 131072
        reasoningEfforts: { off: none, high: high }
        defaultEffort: high
      - id: retired-model            # retire a shipped model
        enabled: false

Field semantics:

  • reasoningEfforts keys are the levels the picker offers (off, minimal, low, medium, high, xhigh, max); values are the spellings sent to Ollama (off: none). false declares a model with no thinking control; omission inherits the built-in entry, and an id neither the catalog nor the entry describes takes the standard ladder (off/low/medium/high/max) with no default claimed.
  • defaultEffort must be one of the offered levels. It is materialized when a session picks none; otherwise the model's own default applies.
  • An override inherits the built-in default level only while it leaves reasoningEfforts alone — changing the level set makes the declaration authoritative.
  • A declared defaultEffort outside the offered set fails at mount instead of degrading silently.

Local Ollama / self-hosted endpoints

- id: llm-ollama-cloud
  config:
    baseURL: http://localhost:11434/api
    apiKeyEnv: ''                     # a local server needs no bearer token

baseURL is normalized: a bare host gains /api, a /v1 spelling maps back to /api, and an explicit custom path (https://gateway.example/ollama) is kept as written.

Web search and fetch

The plugin registers Ollama's /api/web_search and /api/web_fetch as ctx.web providers — registering changes no deployment policy. To use them:

- id: web
  config:
    searchProvider: ollama-cloud
    fetchProvider: ollama-cloud

Both share the route's credential reference and baseURL. Requests carry the credential, so redirects fail closed; each attempt has a 15-second budget and one retry on a transient pre-response transport failure.