Zhangbo-cn/dsh-vision-plugin--packages-vision ↗★ 0
@zhangbo-cn/dsh-vision
DeepSeek Harness 的视觉能力缝隙:一个带有自动选择 describe() 执行功能的提供商注册表,让纯文本模型能够“理解”图像。
AI 分析
核心用途是作为视觉服务定义和注册表,管理图像描述适配器。适合需要为 DeepSeek Harness 接入外部多模态视觉能力的开发者。是构建 vision 插件生态的基础依赖包。
安裝
此插件尚未提供可驗證的 bundle,或相容性檢查未通過。請先閱讀倉庫說明。 閱讀完整 README ↗
說明文件
閱讀完整 README ↗dsh-vision-plugin
Vision capability for DeepSeek Harness: lets a text-only model "understand" an image by routing to an external OpenAI-compatible multimodal API. Built as a standalone dsh-plugin from the official vision capability seam proposal.
Packages
| Package | Role |
|---|---|
@zhangbo-cn/dsh-vision | Service Definition: ctx.vision (registerAdapter, describe, listProviders) |
@zhangbo-cn/dsh-vision-openai-compatible | Provider: OpenAI-compatible chat-completions adapter |
@zhangbo-cn/dsh-tool-vision | Consumer: view_image tool |
Install
pnpm add @zhangbo-cn/dsh-vision @zhangbo-cn/dsh-vision-openai-compatible @zhangbo-cn/dsh-tool-vision
Mount in your cordis.yml:
- id: vision
name: '@zhangbo-cn/dsh-vision'
- id: vision-openai-compatible
name: '@zhangbo-cn/dsh-vision-openai-compatible'
config:
baseURL: 'https://api.example.com/v1' # required at request time
model: 'gpt-4o' # required at request time
apiKeyEnv: 'OPENAI_API_KEY' # env var holding the key
- id: tool-vision
name: '@zhangbo-cn/dsh-tool-vision'
Then ask the model: "use view_image to look at ./screenshot.png" — it reads the file, commits the bytes through the attachment seam, and returns a text description from your configured vision model.
How it works
view_image(file_path, prompt)
→ ctx.fs reads the image bytes
→ attachments.saveImage (durable, content-addressed)
→ ctx.vision.describe({ ref, prompt })
→ vision provider posts a data:image/...;base64 image_url to /chat/completions
→ returns text
Image input reuses the durable ImageAttachmentRef from the attachment seam; output is text (no ImageBlock), so it is independent of whether the harness LLM route itself accepts images.
Requirements
- DeepSeek Harness with an attachment store (
dsh-attachment-local) and filesystem (dsh-fs-local). - A configured OpenAI-compatible multimodal endpoint (any OpenAI-chat-completions-compatible vision model).
Development
npm install
npm run build -ws
npx vitest run
Tests include a real Loader composition booting the plugin with a fake in-memory vision provider (52 tests).
License
MIT