Zhangbo-cn/dsh-vision-plugin--packages-vision-openai-compatible ↗★ 0
@zhangbo-cn/dsh-vision-openai-compatible
OpenAI-compatible chat-completions vision provider for the DeepSeek Harness vision seam (ctx.vision).
AI Analysis
该插件为仅支持文本的模型提供视觉能力,通过路由至外部兼容 OpenAI 的多模态 API 来理解图像。适合需要为 DeepSeek Harness 引入图像识别与分析任务的用户。
Install
This plugin has no verified bundle, or compatibility checks failed. Read the repository notes first. Read the full README ↗
README
Read the full README ↗dsh-vision-plugin
Vision capability for DeepSeek Harness: lets a text-only model "understand" an image by routing to an external OpenAI-compatible multimodal API. Built as a standalone dsh-plugin from the official vision capability seam proposal.
Packages
| Package | Role |
|---|---|
@zhangbo-cn/dsh-vision | Service Definition: ctx.vision (registerAdapter, describe, listProviders) |
@zhangbo-cn/dsh-vision-openai-compatible | Provider: OpenAI-compatible chat-completions adapter |
@zhangbo-cn/dsh-tool-vision | Consumer: view_image tool |
Install
pnpm add @zhangbo-cn/dsh-vision @zhangbo-cn/dsh-vision-openai-compatible @zhangbo-cn/dsh-tool-vision
Mount in your cordis.yml:
- id: vision
name: '@zhangbo-cn/dsh-vision'
- id: vision-openai-compatible
name: '@zhangbo-cn/dsh-vision-openai-compatible'
config:
baseURL: 'https://api.example.com/v1' # required at request time
model: 'gpt-4o' # required at request time
apiKeyEnv: 'OPENAI_API_KEY' # env var holding the key
- id: tool-vision
name: '@zhangbo-cn/dsh-tool-vision'
Then ask the model: "use view_image to look at ./screenshot.png" — it reads the file, commits the bytes through the attachment seam, and returns a text description from your configured vision model.
How it works
view_image(file_path, prompt)
→ ctx.fs reads the image bytes
→ attachments.saveImage (durable, content-addressed)
→ ctx.vision.describe({ ref, prompt })
→ vision provider posts a data:image/...;base64 image_url to /chat/completions
→ returns text
Image input reuses the durable ImageAttachmentRef from the attachment seam; output is text (no ImageBlock), so it is independent of whether the harness LLM route itself accepts images.
Requirements
- DeepSeek Harness with an attachment store (
dsh-attachment-local) and filesystem (dsh-fs-local). - A configured OpenAI-compatible multimodal endpoint (any OpenAI-chat-completions-compatible vision model).
Development
npm install
npm run build -ws
npx vitest run
Tests include a real Loader composition booting the plugin with a fake in-memory vision provider (52 tests).
License
MIT