Zhangbo-cn/dsh-vision-plugin--packages-tool-vision0

@zhangbo-cn/dsh-tool-vision

DeepSeek Harness 视觉缝隙(ctx.vision)上面向模型的 view_image 工具:读取图像文件,持久化提交其字节,并返回来自外部视觉提供商的文本描述。

AI 分析

核心用途是为模型提供 view_image 工具以读取并描述图像。适合需要让纯文本模型理解本地图像文件的任务。必要条件是需安装并配置配套的 vision 服务和外部兼容的 API 提供商。

包名
@zhangbo-cn/dsh-tool-vision
版本
0.1.0
许可证
MIT
最近更新
2026年8月15日

安装

此插件尚未提供可验证的 bundle,或兼容性检查未通过。请先阅读仓库说明。 阅读完整 README ↗

dsh-vision-plugin

Vision capability for DeepSeek Harness: lets a text-only model "understand" an image by routing to an external OpenAI-compatible multimodal API. Built as a standalone dsh-plugin from the official vision capability seam proposal.

Packages

PackageRole
@zhangbo-cn/dsh-visionService Definition: ctx.vision (registerAdapter, describe, listProviders)
@zhangbo-cn/dsh-vision-openai-compatibleProvider: OpenAI-compatible chat-completions adapter
@zhangbo-cn/dsh-tool-visionConsumer: view_image tool

Install

pnpm add @zhangbo-cn/dsh-vision @zhangbo-cn/dsh-vision-openai-compatible @zhangbo-cn/dsh-tool-vision

Mount in your cordis.yml:

- id: vision
  name: '@zhangbo-cn/dsh-vision'

- id: vision-openai-compatible
  name: '@zhangbo-cn/dsh-vision-openai-compatible'
  config:
    baseURL: 'https://api.example.com/v1'   # required at request time
    model: 'gpt-4o'                          # required at request time
    apiKeyEnv: 'OPENAI_API_KEY'              # env var holding the key

- id: tool-vision
  name: '@zhangbo-cn/dsh-tool-vision'

Then ask the model: "use view_image to look at ./screenshot.png" — it reads the file, commits the bytes through the attachment seam, and returns a text description from your configured vision model.

How it works

view_image(file_path, prompt)
  → ctx.fs reads the image bytes
  → attachments.saveImage (durable, content-addressed)
  → ctx.vision.describe({ ref, prompt })
      → vision provider posts a data:image/...;base64 image_url to /chat/completions
      → returns text

Image input reuses the durable ImageAttachmentRef from the attachment seam; output is text (no ImageBlock), so it is independent of whether the harness LLM route itself accepts images.

Requirements

  • DeepSeek Harness with an attachment store (dsh-attachment-local) and filesystem (dsh-fs-local).
  • A configured OpenAI-compatible multimodal endpoint (any OpenAI-chat-completions-compatible vision model).

Development

npm install
npm run build -ws
npx vitest run

Tests include a real Loader composition booting the plugin with a fake in-memory vision provider (52 tests).

License

MIT