Asheblog/dsh-ollama-cloud ↗★ 0
dsh-ollama-cloud
Ollama Cloud provider for DeepSeek Harness: one-click provider setup with per-model reasoning-effort control and account usage in the Models page and sidebar. 适合使用Ollama Cloud模型并需要精细调节推理强度的用户。
설치
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:Asheblog/dsh-ollama-cloudUsage
- Pick an
ollama-cloudmodel in the composer's model menu (or/model). - Pick an effort in the same menu. The offered levels depend on the model, and the default comes from Ollama's own model metadata.
- Thinking streams into the transcript like any other provider. The effort is stored per session; a running step keeps the level it started with.
Shipped models and levels
Snapshot of live metadata (/api/tags + /api/show) taken 2026-09-29:
| Model | Context | Vision | Levels | Default |
|---|---|---|---|---|
deepseek-v4.1-flash | 1,048,576 | ✔ | Off / Low / High / Max | High |
deepseek-v4-pro:0813 | 1,048,576 | — | Off / Low / High / Max | Low |
kimi-k3 | 1,048,576 | ✔ | Off / Low / High / Max | Max |
kimi-k2.6 | 262,144 | ✔ | Off / High | High |
kimi-k2.7-code | 262,144 | ✔ | Off / High | High |
glm-5.3 | 1,048,576 | — | Low / High / Max | Max |
glm-5.3-flash | 1,048,576 | ✔ | Low / High / Max | Max |
glm-5.2 | 1,048,576 | — | Off / High / Max | High |
gpt-oss:120b | 131,072 | — | Low / Medium / High | Medium |
gpt-oss:20b | 131,072 | — | Low / Medium / High | Medium |
minimax-m3 | 512,000 | ✔ | Off / Low / Medium / High / Max | — |
minimax-m2.7 | 196,608 | — | High | High |
nemotron-3-ultra | 262,144 | — | Off / High | High |
nemotron-3-super | 262,144 | — | Off / High | High |
nemotron-3-nano:30b | 262,144 | — | Off / High | High |
gemma4:31b | 262,144 | ✔ | Off / High | Off |
mistral-large-3:675b | 262,144 | ✔ | (no thinking control) | — |
Models whose metadata is a boolean switch (kimi-k2.6, gemma4:31b, …) expose Off and High only — on the wire, High means "thinking on". minimax-m2.7 reports [true], so thinking cannot be switched off and only High is offered. minimax-m3 reports the thinking capability without a level ladder, so it takes the standard ladder and claims no default; the same applies to any model neither the catalog nor your configuration describes.
Refreshing the catalog
Ollama retires cloud models (a retired id answers HTTP 410). The shipped catalog is a snapshot, not an authority:
- Discover from the endpoint: any surface calling
llm/discoverModelslists what the endpoint currently serves, with context windows and input modalities. Only models/api/showcan fully describe become candidates. - Edit by hand: override
modelsin the plugin configuration (below).enabled: falsehides a shipped model.
Cloud usage in the UI
The plugin ships a browser half that renders in the host's own surfaces — no separate page, and no provider-UI shell plugin needed:
- Models page card: inside the Ollama Cloud row (
Settings → Models), one meter per billing window with its remaining share and reset line, the primary window's per-model request counts, the write-only API key field, and a refresh button. - Sidebar row: a compact remaining-quota line under the session list; click it for the per-window detail. It refreshes when the sidebar mounts and every 15 minutes after that.
The card reads usage through this plugin's own host channel (/ollama-cloud),
so the API key stays on the host and never reaches the browser. A local or
self-hosted endpoint answers 404 on /usage; that renders as "this endpoint
does not report cloud usage" rather than an error, and the last good snapshot
keeps showing instead of disappearing.
The usage channel needs the composition's
connectionrow to injectwebServer; this bundle's patch adds that (profiles without a connection row skip it). After installing or updating the plugin, restart the harness once — until then the card says so itself.
Cloud usage in the UI
The plugin ships a browser half that renders in the host's own surfaces — no separate page, and no provider-UI shell plugin needed:
- Models page card: inside the Ollama Cloud row (
Settings → Models), one meter per billing window with its remaining share and reset line, the primary window's per-model request counts, the write-only API key field, and a refresh button. - Sidebar row: a compact remaining-quota line under the session list; click it for the per-window detail. It refreshes when the sidebar mounts and every 15 minutes after that.
The card reads usage through this plugin's own host channel (/ollama-cloud),
so the API key stays on the host and never reaches the browser. A local or
self-hosted endpoint answers 404 on /usage; that renders as "this endpoint
does not report cloud usage" rather than an error, and the last good snapshot
keeps showing instead of disappearing.
The usage channel needs the composition's
connectionrow to injectwebServer; this bundle's patch adds that (profiles without a connection row skip it). After installing or updating the plugin, restart the harness once — until then the card says so itself.
Configuration
Every field is editable from the plugin settings page and overridable per row in the profile's cordis.patch.yml:
- id: llm-ollama-cloud
name: 'dsh-ollama-cloud'
config:
apiKeyEnv: OLLAMA_API_KEY # credential reference; empty string = no auth (local servers)
baseURL: https://ollama.com/api # native API root; local Ollama: http://localhost:11434/api
maxTokens: 32768 # output cap for models that declare none
defaultContextWindow: 262144 # context fallback for models that declare none
streamIdleTimeoutMs: 300000 # maximum idle time between stream reads
requestTimeoutMs: 15000 # per-attempt budget for the non-chat requests (discovery, web search/fetch)
retryPolicy: # executed by dsh-llm-retry
mode: normal
maxRetries: 5
models: # merges over the built-in catalog by id; new ids append
- id: gpt-oss:20b
contextWindow: 131072
reasoningEfforts: { off: none, low: low, medium: medium, high: high }
defaultEffort: medium
- id: brand-new-model # a model newer than the snapshot
name: Brand New
contextWindow: 131072
reasoningEfforts: { off: none, high: high }
defaultEffort: high
- id: retired-model # retire a shipped model
enabled: false
Field semantics:
reasoningEffortskeys are the levels the picker offers (off,minimal,low,medium,high,xhigh,max); values are the spellings sent to Ollama (off: none).falsedeclares a model with no thinking control; omission inherits the built-in entry, and an id neither the catalog nor the entry describes takes the standard ladder (off/low/medium/high/max) with no default claimed.defaultEffortmust be one of the offered levels. It is materialized when a session picks none; otherwise the model's own default applies.- An override inherits the built-in default level only while it leaves
reasoningEffortsalone — changing the level set makes the declaration authoritative. - A declared
defaultEffortoutside the offered set fails at mount instead of degrading silently.
Local Ollama / self-hosted endpoints
- id: llm-ollama-cloud
config:
baseURL: http://localhost:11434/api
apiKeyEnv: '' # a local server needs no bearer token
baseURL is normalized: a bare host gains /api, a /v1 spelling maps back to /api, and an explicit custom path (https://gateway.example/ollama) is kept as written.
Web search and fetch
The plugin registers Ollama's /api/web_search and /api/web_fetch as ctx.web providers — registering changes no deployment policy. To use them:
- id: web
config:
searchProvider: ollama-cloud
fetchProvider: ollama-cloud
Both share the route's credential reference and baseURL. Requests carry the credential, so redirects fail closed; each attempt has a 15-second budget and one retry on a transient pre-response transport failure.