tiefeiyu/dsh-see-image ↗★ 1

dsh-see-image

把图片转给兼容 OpenAI 的视觉模型,再将文字描述回传给纯文本模型。 适合纯文本模型需读图的场景;需配置视觉模型端点与密钥。

套件
dsh-see-image
相容性
待驗證
Harness 依賴範圍
*
版本
1.0.0
授權
MIT
最近更新
2026年8月14日

安裝

此插件尚未提供可驗證的 bundle,或相容性檢查未通過。請先閱讀倉庫說明。 閱讀完整 README ↗

Configuration

KeyDefaultDescription
baseURLhttps://api.individual.githubcopilot.comOpenAI-compatible endpoint; the plugin appends /chat/completions
modelgpt-4.1Vision model ID
apiKeyEnvVISION_API_KEYEnv var name holding the API key for non-Copilot backends; empty sends no Authorization header (keyless local endpoints like Ollama)
maxTokens1024Max output tokens (Zhipu glm-4v-flash caps at 1024; raise it for larger models)
timeoutMs90000Request timeout
maxBytes15728640Max image size (15 MB)
prompt(detailed Chinese description instruction)Default question; a question argument passed at call time takes precedence

Usage

In conversation, just say "look at this image / read this screenshot" — the model locates the file and calls the tool itself. You can also state the question explicitly:

see_image(file_path="C:\\Users\\me\\Desktop\\error.png", question="What is the full text of this error?")

question may be omitted; the default output covers: scene → verbatim text transcription → color/shape/layout details → explanation of charts/UI/errors (in Chinese).