Zhangbo-cn/dsh-vision-plugin--packages-vision ↗★ 0
@zhangbo-cn/dsh-vision
Vision capability seam for DeepSeek Harness: a provider registry with auto-selecting describe() execution, letting a text-only model 'understand' an image.
AI Analysis
核心用途是作为视觉服务定义和注册表,管理图像描述适配器。适合需要为 DeepSeek Harness 接入外部多模态视觉能力的开发者。是构建 vision 插件生态的基础依赖包。
Install
This plugin has no verified bundle, or compatibility checks failed. Read the repository notes first. Read the full README ↗
README
Read the full README ↗dsh-vision-plugin
Vision capability for DeepSeek Harness: lets a text-only model "understand" an image by routing to an external OpenAI-compatible multimodal API. Built as a standalone dsh-plugin from the official vision capability seam proposal.
Packages
| Package | Role |
|---|---|
@zhangbo-cn/dsh-vision | Service Definition: ctx.vision (registerAdapter, describe, listProviders) |
@zhangbo-cn/dsh-vision-openai-compatible | Provider: OpenAI-compatible chat-completions adapter |
@zhangbo-cn/dsh-tool-vision | Consumer: view_image tool |
Install
pnpm add @zhangbo-cn/dsh-vision @zhangbo-cn/dsh-vision-openai-compatible @zhangbo-cn/dsh-tool-vision
Mount in your cordis.yml:
- id: vision
name: '@zhangbo-cn/dsh-vision'
- id: vision-openai-compatible
name: '@zhangbo-cn/dsh-vision-openai-compatible'
config:
baseURL: 'https://api.example.com/v1' # required at request time
model: 'gpt-4o' # required at request time
apiKeyEnv: 'OPENAI_API_KEY' # env var holding the key
- id: tool-vision
name: '@zhangbo-cn/dsh-tool-vision'
Then ask the model: "use view_image to look at ./screenshot.png" — it reads the file, commits the bytes through the attachment seam, and returns a text description from your configured vision model.
How it works
view_image(file_path, prompt)
→ ctx.fs reads the image bytes
→ attachments.saveImage (durable, content-addressed)
→ ctx.vision.describe({ ref, prompt })
→ vision provider posts a data:image/...;base64 image_url to /chat/completions
→ returns text
Image input reuses the durable ImageAttachmentRef from the attachment seam; output is text (no ImageBlock), so it is independent of whether the harness LLM route itself accepts images.
Requirements
- DeepSeek Harness with an attachment store (
dsh-attachment-local) and filesystem (dsh-fs-local). - A configured OpenAI-compatible multimodal endpoint (any OpenAI-chat-completions-compatible vision model).
Development
npm install
npm run build -ws
npx vitest run
Tests include a real Loader composition booting the plugin with a fake in-memory vision provider (52 tests).
License
MIT