返回目录

dsh-vision

编辑精选维护状态: 活跃

william-jin-cmu/dsh-vision

dsh 插件:给纯文本 DeepSeek 加视觉——view_image 工具桥接任意 OpenAI 兼容 VLM(默认智谱免费档,实测 4 厂商 10 模型)

前往 GitHub
$ git clone https://github.com/dsh-external/dsh-vision ~/dsh-plugins/dsh-vision

30

星标

5

Fork

TypeScript

语言

BSD-3-Clause

许可证

2026-08-05

创建于

2026-08-13

最近推送

原生挂载到 dsh 的视觉插件:注册 view_image 工具,把图片和问题转发给任意 OpenAI 兼容 VLM,答案以文本返回,所有入口同时获得视觉。

DSH 适配

原生运行时

作者声明

安全审计

未审计

最后核验

2026-09-03

许可证

BSD-3-Clause

01它能帮你完成什么?

  • Let text-only DeepSeek answer visual questions about images (OCR, counting, chart reading, UI layout)

    A `view_image` tool the model calls with an image path and a question; returns the answer as text

    DeepSeek Harness (dsh) users who need vision on a text-only model across web / TUI / remote channels

  • Route image understanding to any OpenAI-compatible VLM backend (free Zhipu, Qwen, Doubao, local Ollama)

    Configurable baseURL + apiKey + model; automatic free-tier fallback chain keeps zero-config answers working

    Users who want to keep their own BYOK / local VLM endpoint instead of a paid hosted vision service

02如何接入 DeepSeek Harness?

前置条件

  • DeepSeek Harness (dsh) installed and on PATH
  • A VLM endpoint (default uses free Zhipu glm-4.6v-flash; other backends need an API key or a local endpoint)

安装步骤

  1. 01

    Clone the plugin: `git clone https://github.com/dsh-external/dsh-vision ~/dsh-plugins/dsh-vision`

    $ git clone https://github.com/dsh-external/dsh-vision ~/dsh-plugins/dsh-vision

  2. 02

    Symlink host dependencies (@deepseek-ai/dsh-tools and schemastery) into the plugin's node_modules

  3. 03

    Append the mount block to ~/.dsh/config.yaml and restart dsh

验证接入成功

  • Ask dsh about an image, e.g. "what error is in ~/Desktop/error.png"; the model should call view_image and return text

03DSH 适配与能力边界

DSH 适配原生运行时

以原生方式挂载进 dsh 个人配置覆盖层(~/.dsh/config.yaml),不依赖任何第三方管理器,以原生 cordis 插件形态运行。

  • view_image tool

    image path / URL + natural-language question→text answer returned to the model

  • Multi-backend routing

    configured baseURL + apiKey + model→forwards image + question to the VLM endpoint, returns text

    sends image data to an external OpenAI-compatible VLM endpoint (or a local endpoint)
  • Free-tier fallback chain

    rate-limited default model (HTTP 429)→auto-downgrades glm-4.6v-flash → glm-4.1v-thinking-flash → glm-4v-flash

  • Zero-dependency cordis bridge

    OpenAI-compatible /chat/completions + image_url→native TypeScript cordis plugin, no Python / uv / MCP

    strips <think> reasoning blocks from thinking modelsauto-redacts API key in error messages

04适合谁?何时不该用?

适合

  • DeepSeek Harness (dsh) users who need vision on a text-only model across web / TUI / remote channels
  • Users who want to keep their own BYOK / local VLM endpoint instead of a paid hosted vision service

不适合

  • Before marisa#2 is fixed, you must manually symlink the host dsh node_modules (@deepseek-ai/dsh-tools and schemastery) or the plugin fails to load.

05兼容性、维护与安全提示

  • Images are forwarded to an external OpenAI-compatible VLM endpoint, so the plugin needs network access and (for non-local backends) an API key. Localhost endpoints need no key.
  • Before marisa#2 is fixed, you must manually symlink the host dsh node_modules (@deepseek-ai/dsh-tools and schemastery) or the plugin fails to load.
  • Zhipu's free model is rate-limited (429) on a shared pool; the plugin auto-downgrades across three free models to keep zero-config answers working.
2026-08-052026-08-13作者未说明

BSD-3-Clause · no releases yet (curated; latest push 2026-08-13)

06常见问题

dsh-vision 是做什么的?

它注册一个 view_image 工具,让纯文本的 DeepSeek 也能看图——OCR、数数、读图表、看 UI 布局等任意视觉问题——把图片和问题转发给任意 OpenAI 兼容的 VLM,答案以文本返回。

怎么安装?

默认是原生挂载:克隆仓库,把宿主 dsh 的 node_modules(@deepseek-ai/dsh-tools 和 schemastery)软链进插件,再把挂载块追加到 ~/.dsh/config.yaml 并重启 dsh。DSH Companion 已随应用自带、零安装。

支持哪些视觉后端?

任意 OpenAI 兼容 VLM:智谱 glm-4.6v-flash(默认免费)、通义 qwen vl 系列、火山豆包、本地 Ollama 等——一套 baseURL+apiKey+model 通吃。本地端点无需 key。

免费档被限流怎么办?

默认配置会自动依次降级 glm-4.6v-flash → glm-4.1v-thinking-flash → glm-4v-flash,保证零配置也总能出答案;可用 fallbackModels 覆盖。

modlens

作者 liustack

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

精选设计、媒体与视觉TypeScript
4,072124

agent-vision-toolkit

作者 anionex

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

设计、媒体与视觉Python
1,21748

dsh-vision-router

作者 ysr666

为纯文本 DeepSeek Harness 智能体提供「视觉」能力,内置免密钥视觉链路与像素级视觉工具,一条命令安装,无需 Python。

精选设计、媒体与视觉JavaScript
1,12851

dsh-vision-toolkit

作者 anionex

[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

精选设计、媒体与视觉TypeScript
88447

08数据与来源

  • 作者声明github.com72978aa176df…

    给纯文本的 DeepSeek 加上眼睛。Vision for text-only DeepSeek.

  • 作者声明github.com72978aa176df…

    本插件注册一个 `view_image` 工具:模型带着问题调用它(OCR、数数、读图表、看 UI 布局……任意视觉问题)

页面基于项目公开文档、仓库元数据和 DSH Plugins 的结构化解析生成;最后核验于 2026-09-03。发现错误?提交更正。

🏆

最佳 DeepSeek Harness 插件

从全目录挑出的 12 个值得优先安装的插件,覆盖各个分类。

DSH Plugins 是独立的 DeepSeek Harness 插件市场,与 DeepSeek 官方无关,也不代表官方背书。第三方插件未经安全审计,安装前请审查源码。

每周获取最新的 DeepSeek Harness 插件,绝不滥发。