zzdream67/dsh-vision-bridge ↗★ 0
@zzdream67/dsh-vision-bridge
拦截模型请求,用视觉模型描述替换图片,让纯文本模型读图。 适合使用纯文本模型但需处理图片的用户,需配置可用的视觉模型端点。
安装
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:zzdream67/dsh-vision-bridge说明文档
阅读完整 README ↗Configuration
The plugin ships a browser-side bundle (dist/client.js) that registers its own section into the Web settings page's settings.section slot. Every field is editable there, applies live, and needs no restart.
Three buttons in that page are worth knowing about:
- Test and enable sends one 16×16 PNG to the selected vision model. The declaration is written only if the endpoint accepts the image, and is withdrawn again on failure. Capability is tested rather than assumed, because an OpenAI-compatible
/v1/modelsreply carries no modality information. - Enable directly skips the test and writes the declaration. Use it when you already know the model can see.
- Reload re-reads the configuration from disk. The desktop client has no reload gesture, so this is the only way to pick up a change made in the Models page, in another window, or by hand. It asks first if you have unsaved edits.
Host-side schema registration alone does not render a form. The section exists because this plugin authors that browser bundle by hand; the build preset for DSH's client format is not published, but the format itself is plain and stable.
Why a declaration is required first
The admission gate (dsh-host-apiproxy) checks the model's declaration before the image enters the session:
if (modelInfo.inputModalities !== undefined && !modelInfo.inputModalities.includes('image'))
return err(request, { details: { reason: 'MODEL_DOES_NOT_SUPPORT_IMAGES' } })
A text-only model must declare input: [text, image], or the image never arrives and this plugin's conversion never runs.
manageDeclarations (default: on) writes that declaration for each model in bridge and withdraws it when you turn the switch off, disable the plugin, or unload it. With the plugin loaded and the model bridged, the route genuinely does accept image input, because conversion happens before the provider is called.
input: [text, image] is an admission label. Every place the host reads it is guarded by "does this request contain an image", so a text-only request behaves identically whether or not the declaration is present. It does not affect token accounting, context window, sampling, or model behaviour.
Declaring a model without bridging it moves the failure later: the host admits the image and the provider then rejects it. Declarations therefore follow the bridge list rather than being applied broadly.
Fields
| Field | Default | Meaning |
|---|---|---|
enabled | true | Master switch. When off the plugin stays loaded but intervenes in nothing and withdraws its declarations, like uninstalling without touching node_modules |
visionProvider | '' | Vision route id. Empty leaves the plugin inert (see Behavior details) |
visionModel | '' | Vision model id in that route. Empty leaves the plugin inert (see Behavior details) |
prompt | see below | Caption instruction sent with each image |
bridge | [] | Text-only models to bridge |
manageDeclarations | true | Let the plugin write/withdraw modality declarations |
cacheSize | 64 | Cached transcriptions; 0 disables caching |
timeoutMs | 120000 | Per-image caption timeout |
verbose | false | One log line per conversion (never logs transcription text) |
Or in cordis.patch.yml:
- id: zz-vision-bridge
name: '@zzdream67/dsh-vision-bridge'
config:
visionProvider: my-local-route
visionModel: my-vision-model
bridge:
- provider: my-text-route
model: my-text-model
manageDeclarations: true
A model not listed in bridge is untouched and keeps the host's exact behavior.
Models under the built-in deepseek-official route cannot be bridged: its modality is hardcoded in the plugin that provides it, and its adapter refuses image content. To bridge DeepSeek, add a custom model provider with a route using openai-completions at https://api.deepseek.com/v1, then bridge the models under that route.
The default prompt
It asks for transcription, not interpretation: transcribe all visible text faithfully, describe layout and structure, never speculate or fill gaps, and state plainly what is illegible. The reasoning belongs to the model consuming the transcription; a vision model that editorializes costs the caller information it can never recover.