catalog / media-vision
Design, Media & Vision
Every Design, Media & Vision plugin in our curated awesome list, synced from GitHub.
Multimodal capabilities for the harness: vision toolkits, OCR, image generation, design canvases and media processing. They give the agent eyes and a sketchpad, registering as tool providers the model can call mid-task.
Synced from our GitHub list · Catalog updated 2026-09-16
509 plugins
modlens
by liustack
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
agent-vision-toolkit
by anionex
Vision toolkit and skills that give text-only LLMs eyes — multi-image understanding, image Q&A, OCR, frontend UI restoration and GUI automation, with optional agent integration.
dsh-vision-router
by ysr666
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
dsh-vision-toolkit
by anionex
为纯文本 DSH Agent 提供 10 个结构化视觉工具:意图感知图片问答、长截图 OCR、原始像素 grounding、UI 还原、像素 diff 等
dsh-image-gen
by shanliuling
AI image studio for DeepSeek Harness — generate, edit & compare images in chat, with 500+ prompts, gallery, multi-model workflows and ComfyUI.
dsh-openpencil
by zseven-w
The DeepSeek Harness plugin for OpenPencil — preview, inspect, and edit real .op documents inside a conversation.
dsh-video-lens
by dundunhan
Two tools that let text-only DSH agents understand local video: frame extraction + vision captioning and ASR transcription, provider-configurable
dsh-vision
by oil-oil
Near-native image understanding for DeepSeek Harness
dsh-imagegen
by dickpy
DSH (DeepSeek Harness) Web GUI AI image generation plugin: text-to-image & image-to-image via OpenAI-compatible endpoints (gpt-image-2), with shared cross-device history.
dsh-comfyui
by fandc520
Agent tools comfyui_run/comfyui_object_info/comfyui_workflow submit and inspect ComfyUI workflows, render results in-chat.
patentradar
by yuc16
Patent infringement analyzer that turns a patent publication number into a competitor infringement report — also ships as a skill any agent (Codex, Claude Code) can call.
dsh-design-qa
by sunxin-ai
Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.
dsh-vision-complete
by yts1919
Windows skill pack: SKILL.md + vision.py (image/OCR/object-detect/video/voice/PDF via cloud Qwen), registers qwen-mm-plugins MCP into cordis.patch.yml, clipboard screenshot tool.
picturereader
by jing-hy
DSH plugin: pixel-to-text image reading for text-only models. image_scan/image_ocr/image_sample tools + image-reading skill (34-image trained methodology). Pure local, optional PaddleOCR.
dsh-vision
by william-jin-cmu
Registers a view_image tool that bridges text-only DeepSeek to any OpenAI-compatible VLM endpoint for OCR, counting, chart reading and UI-layout questions.
dsh-design-mode
by kaichencurry
Visual execution layer for DeepSeek Harness—guided intent, infinite canvas, contextual image tools, and chat-native comments.
dsh-media-skills
by mjorgin
Free image reading & generation for DeepSeek Harness (rc.7 / rc.8 / v0.1.1-rc.1 / rc.2 / v0.1.2-alpha.3) — paste-image reading with auto vision transcription, DeepSeek-V4-Flash-Vision-Exp / GLM-4V-Flash / SenseNova / Gemini failover, Kolors + U1 Fast generation. No keys in repo.
deepseek-visionary
by xlight
给 DSH 接入 DeepSeek 网页版原生多模态视觉:deepseek_vision/status/login/logout 4 个宿主级原生工具,宿主进程内 spawn visionary-server(Rust 单二进制:PoW→上传→fork→HIF 签名→SSE 流式),支持 CDP 浏览器自动登录;另有 vision CLI 与内嵌 skill。
dsh-plugin-aigc-canvas
by huanlinoto
Provider-agnostic AIGC HTTP bridge + infinite canvas + ffmpeg post-processing (aigc_http_request, aigc_canvas_place, aigc_media_edit).
dsh-visual-plugin
by jyh20030112
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
dsh-vision-proxy
by flyvhidbwo
DeepSeek brain + automatic image transcription — proxies attached images to a VLM (DashScope qwen by default) and feeds the transcribed text back to DeepSeek.
dsh-vision-opencode
by poiuyjie
DSH plugin: Auto-convert images to text for pure-text LLMs (DeepSeek etc.) via any vision model. No need to switch your main model.
dsh-ocr-plugin
by crazy222123
Local OCR provider (RapidOCR fast + DeepSeek-OCR-2/llama.cpp deep) registered as 'ocr' service on the llm-deepseek adapter seam; converts images to text blocks before API send.
dsh-vision
by linenxi-ctrl
为 DeepSeek Harness 增加外挂识图模型:圆形鲸鱼按钮配置面板、发送图片识图自动回传当前会话、为 agent 注入 screenshot/recognize_image 工具、多协议自动适配(可配地址/密钥/模型/提示词/代理)。
dsh-bilibili
by czx2244
Bilibili video analysis: metadata, transcript (ASR fallback via Bijian/sherpa-onnx/whisper.cpp), comments, danmaku, and sharp keyframes with optional local vision descriptions.
dsh-mmx-bridge
by welsione
MiniMax multimodal bridge: one `mmx_bridge` tool covers image understanding/generation, video, TTS, music, cover, web search and quota; optional `web_search`/`read_image` takeover; inline players/image previews right in the Web GUI (npm: `dsh-mmx-bridge`).
dsh-chat-imagine
by corrinehu
Inline image generation with bundled cli-image-gen recovery skill; analyze_image reads local/URL images into structured JSON evidence; set_image_default for channel selection.
dsh-image2-draw
by junelearn
Image2 生图插件
dsh-mermaid
by aks1st
在 DSH 中渲染 Mermaid 图表,将图表代码转为可视化。
dsh-windows-ocr
by maxwell-feng
dsh plugin: OCR attached images locally with the built-in Windows OCR engine — text-only models can see, privacy-first
