A native dsh plugin that adds vision to text-only DeepSeek by routing images to an OpenAI-compatible VLM and returning answers as text.
DSH integration
Native runtime
Author-claimed
Safety audit
Unaudited
Last verified
2026-09-03
License
BSD-3-Clause
01What can it help you accomplish?
Let text-only DeepSeek answer visual questions about images (OCR, counting, chart reading, UI layout)
A `view_image` tool the model calls with an image path and a question; returns the answer as text
DeepSeek Harness (dsh) users who need vision on a text-only model across web / TUI / remote channels
Route image understanding to any OpenAI-compatible VLM backend (free Zhipu, Qwen, Doubao, local Ollama)
Configurable baseURL + apiKey + model; automatic free-tier fallback chain keeps zero-config answers working
Users who want to keep their own BYOK / local VLM endpoint instead of a paid hosted vision service
02How to install into DeepSeek Harness
Prerequisites
- DeepSeek Harness (dsh) installed and on PATH
- A VLM endpoint (default uses free Zhipu glm-4.6v-flash; other backends need an API key or a local endpoint)
Installation steps
- 01
Clone the plugin: `git clone https://github.com/dsh-external/dsh-vision ~/dsh-plugins/dsh-vision`
$ git clone https://github.com/dsh-external/dsh-vision ~/dsh-plugins/dsh-vision
- 02
Symlink host dependencies (@deepseek-ai/dsh-tools and schemastery) into the plugin's node_modules
- 03
Append the mount block to ~/.dsh/config.yaml and restart dsh
Verify the integration
- Ask dsh about an image, e.g. "what error is in ~/Desktop/error.png"; the model should call view_image and return text
03DSH integration and capability boundaries
Native mount into dsh's personal config overlay (~/.dsh/config.yaml); zero third-party manager, runs as a native cordis plugin.
view_image tool
image path / URL + natural-language question→text answer returned to the model
Multi-backend routing
configured baseURL + apiKey + model→forwards image + question to the VLM endpoint, returns text
sends image data to an external OpenAI-compatible VLM endpoint (or a local endpoint)Free-tier fallback chain
rate-limited default model (HTTP 429)→auto-downgrades glm-4.6v-flash → glm-4.1v-thinking-flash → glm-4v-flash
Zero-dependency cordis bridge
OpenAI-compatible /chat/completions + image_url→native TypeScript cordis plugin, no Python / uv / MCP
strips <think> reasoning blocks from thinking modelsauto-redacts API key in error messages
04Who is it for? When not to use it?
Good for
- DeepSeek Harness (dsh) users who need vision on a text-only model across web / TUI / remote channels
- Users who want to keep their own BYOK / local VLM endpoint instead of a paid hosted vision service
Not for
- Before marisa#2 is fixed, you must manually symlink the host dsh node_modules (@deepseek-ai/dsh-tools and schemastery) or the plugin fails to load.
05Compatibility, maintenance and safety notes
- Images are forwarded to an external OpenAI-compatible VLM endpoint, so the plugin needs network access and (for non-local backends) an API key. Localhost endpoints need no key.
- Before marisa#2 is fixed, you must manually symlink the host dsh node_modules (@deepseek-ai/dsh-tools and schemastery) or the plugin fails to load.
- Zhipu's free model is rate-limited (429) on a shared pool; the plugin auto-downgrades across three free models to keep zero-config answers working.
BSD-3-Clause · no releases yet (curated; latest push 2026-08-13)
06Frequently asked questions
What does dsh-vision do?
It registers a view_image tool that gives text-only DeepSeek the ability to look at images — OCR, counting, chart reading, UI-layout questions — by forwarding the image and your question to an OpenAI-compatible VLM and returning the answer as text.
How do I install it?
Native mount is the default: clone the repo, symlink the host dsh node_modules (@deepseek-ai/dsh-tools and schemastery), and append a mount block to ~/.dsh/config.yaml, then restart dsh. DSH Companion ships it zero-install.
Which vision backends are supported?
Any OpenAI-compatible VLM: Zhipu glm-4.6v-flash (free default), Qwen vl-flash/plus/max, Doubao Seed, local Ollama, and more — one baseURL + apiKey + model covers them all. Local endpoints need no key.
What if the free model is rate-limited?
The default config auto-downgrades glm-4.6v-flash → glm-4.1v-thinking-flash → glm-4v-flash so zero-config setups still get an answer; you can override with fallbackModels.
07Related DSH workflows
modlens
by liustack
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
agent-vision-toolkit
by anionex
Vision toolkit and skills that give text-only LLMs eyes — multi-image understanding, image Q&A, OCR, frontend UI restoration and GUI automation, with optional agent integration.
dsh-vision-router
by ysr666
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
dsh-vision-toolkit
by anionex
为纯文本 DSH Agent 提供 10 个结构化视觉工具:意图感知图片问答、长截图 OCR、原始像素 grounding、UI 还原、像素 diff 等
08Data and sources
给纯文本的 DeepSeek 加上眼睛。Vision for text-only DeepSeek.
本插件注册一个 `view_image` 工具:模型带着问题调用它(OCR、数数、读图表、看 UI 布局……任意视觉问题)
This page is generated from the project’s public documentation, repository metadata and a structured parse of DSH Plugins; last verified on 2026-09-03. Found an error? Submit a correction.
Best DeepSeek Harness Plugins
Twelve plugins worth installing first — picked from the whole catalog, across every category.
