Elohia/dsh-plugin-mm-vision ↗★ 1
dsh-plugin-mm-vision
mm-vision (通感编码器) for DeepSeek Harness — give any text-only LLM the ability to see images via structured spatial text encoding. Registers the mm_vision tool. 适合需让纯文本模型理解图片的用户;需配置视觉API密钥与模型。
インストール
$
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:Elohia/dsh-plugin-mm-visionドキュメント
README 全文を読む ↗⚙️ 配置
配置解析顺序(首个命中):
cordis.patch.yml中本行config字段(安装后可在 profile 的cordis.patch.yml覆盖)- 环境变量 / 配置文件(与原版一致):
export MM_VISION_API_KEY=sk-xxx # 或 DASHSCOPE_API_KEY / QWEN_API_KEY / OPENAI_API_KEY / GEMINI_API_KEY
export MM_VISION_MODEL=qwen-vl-max # 可选,默认 qwen-vl-max
export MM_VISION_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1 # 可选
或写 ~/.config/mm-vision/config.json / 项目根 vision-config.json:
{
"model": "qwen-vl-max",
"baseUrl": "https://dashscope.aliyuncs.com/compatible-mode/v1",
"maxTokens": 2048,
"mode": "auto",
"cacheTTL": 600,
"cacheMax": 100,
"dotMatrix": false
}
用别的视觉模型?改
model+baseUrl即可(OpenAI 兼容协议)。
🎯 使用
安装后在对话中直接让模型看图:
分析这张K线图 F:/data/kline.png
帮我看看 examples/chart.png 里按钮的位置
扫描这张图片并重建像素网格 examples/photo.png
模型会自动调用 mm_vision 工具,返回结构化通感编码:
【图片通感编码(mm-vision)】模式:coords · 模型:qwen-vl-max
1. 【画布】16:9,浅色背景 #f5f6fa
2. 【元素】[矩形 | (10%,10%) | 35%x20% | #4a90d9 | "登录按钮"]…
3. 【关系】…
4. 【图表专用】坐标轴 0-100,转折点 (30%,45%)=52 …