dsh-see-image
把图片转给兼容 OpenAI 的视觉模型,再将文字描述回传给纯文本模型。 适合纯文本模型需读图的场景;需配置视觉模型端点与密钥。
安装
此插件尚未提供可验证的 bundle,或兼容性检查未通过。请先阅读仓库说明。 阅读完整 README ↗
说明文档
阅读完整 README ↗Configuration
| Key | Default | Description |
|---|---|---|
baseURL | https://api.individual.githubcopilot.com | OpenAI-compatible endpoint; the plugin appends /chat/completions |
model | gpt-4.1 | Vision model ID |
apiKeyEnv | VISION_API_KEY | Env var name holding the API key for non-Copilot backends; empty sends no Authorization header (keyless local endpoints like Ollama) |
maxTokens | 1024 | Max output tokens (Zhipu glm-4v-flash caps at 1024; raise it for larger models) |
timeoutMs | 90000 | Request timeout |
maxBytes | 15728640 | Max image size (15 MB) |
prompt | (detailed Chinese description instruction) | Default question; a question argument passed at call time takes precedence |
Usage
In conversation, just say "look at this image / read this screenshot" — the model locates the file and calls the tool itself. You can also state the question explicitly:
see_image(file_path="C:\\Users\\me\\Desktop\\error.png", question="What is the full text of this error?")
question may be omitted; the default output covers: scene → verbatim text transcription → color/shape/layout details → explanation of charts/UI/errors (in Chinese).