catalog / media-vision
Design, Media & Vision
Every Design, Media & Vision plugin in our curated awesome list, synced from GitHub.
Multimodal capabilities for the harness: vision toolkits, OCR, image generation, design canvases and media processing. They give the agent eyes and a sketchpad, registering as tool providers the model can call mid-task.
Synced from our GitHub list · Catalog updated 2026-09-30
527 plugins
modlens
by liustack
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
agent-vision-toolkit
by anionex
Vision toolkit and skills that give text-only LLMs eyes — multi-image understanding, image Q&A, OCR, frontend UI restoration and GUI automation, with optional agent integration.
dsh-vision-router
by ysr666
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
dsh-vision-toolkit
by anionex
为纯文本 DSH Agent 提供 10 个结构化视觉工具:意图感知图片问答、长截图 OCR、原始像素 grounding、UI 还原、像素 diff 等
dsh-image-gen
by shanliuling
AI image studio for DeepSeek Harness — generate, edit & compare images in chat, with 500+ prompts, gallery, multi-model workflows and ComfyUI.
dsh-openpencil
by zseven-w
The DeepSeek Harness plugin for OpenPencil — preview, inspect, and edit real .op documents inside a conversation.
dsh-video-lens
by dundunhan
Two tools that let text-only DSH agents understand local video: frame extraction + vision captioning and ASR transcription, provider-configurable
brewreel
by finderchangchang
Turns a storyboard JSON into validated, compliant vertical promo videos through DSH tools.
scientificfigurelibrary
by xuzhougeng
Local-first MCP App for scientific figures. Import, review, and publish a global library on disk; reuse exact templates in Pi, DeepSeek Harness (dsh), Claude, Codex, Cursor, and Wisp.
dsh-imagegen
by dickpy
DSH (DeepSeek Harness) Web GUI AI image generation plugin: text-to-image & image-to-image via OpenAI-compatible endpoints (gpt-image-2), with shared cross-device history.
dsh-vision
by oil-oil
Near-native image understanding for DeepSeek Harness
dsh-comfyui
by fandc520
Agent tools comfyui_run/comfyui_object_info/comfyui_workflow submit and inspect ComfyUI workflows, render results in-chat.
deepseek-harness-video-director
by chiphoton
🎬AI-Powered Director: MiniMax-H3 Video Generation Plugin Directed by DeepSeek-Harness. 🤖More than Prompting. 🪄Canvas UI.
patentradar
by yuc16
Patent infringement analyzer that turns a patent publication number into a competitor infringement report — also ships as a skill any agent (Codex, Claude Code) can call.
dsh-vision-complete
by yts1919
Windows skill pack: SKILL.md + vision.py (image/OCR/object-detect/video/voice/PDF via cloud Qwen), registers qwen-mm-plugins MCP into cordis.patch.yml, clipboard screenshot tool.
dsh-design-qa
by sunxin-ai
Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.
picturereader
by jing-hy
DSH plugin: pixel-to-text image reading for text-only models. image_scan/image_ocr/image_sample tools + image-reading skill (34-image trained methodology). Pure local, optional PaddleOCR.
dsh-design-mode
by kaichencurry
Visual execution layer for DeepSeek Harness—guided intent, infinite canvas, contextual image tools, and chat-native comments.
dsh-vision
by william-jin-cmu
Registers a view_image tool that bridges text-only DeepSeek to any OpenAI-compatible VLM endpoint for OCR, counting, chart reading and UI-layout questions.
dsh-game-material-master
by universe-st
Generates game sprites, images, and sequence frames through Seedream/MiniMax with local ffmpeg keying, extraction, and compositing.
dsh-media-skills
by mjorgin
Free image reading & generation for DeepSeek Harness (rc.7 / rc.8 / v0.1.1-rc.1 / rc.2 / v0.1.2-alpha.3) — paste-image reading with auto vision transcription, DeepSeek-V4-Flash-Vision-Exp / GLM-4V-Flash / SenseNova / Gemini failover, Kolors + U1 Fast generation. No keys in repo.
deepseek-visionary
by xlight
给 DSH 接入 DeepSeek 网页版原生多模态视觉:deepseek_vision/status/login/logout 4 个宿主级原生工具,宿主进程内 spawn visionary-server(Rust 单二进制:PoW→上传→fork→HIF 签名→SSE 流式),支持 CDP 浏览器自动登录;另有 vision CLI 与内嵌 skill。
dsh-plugin-aigc-canvas
by huanlinoto
Provider-agnostic AIGC HTTP bridge + infinite canvas + ffmpeg post-processing (aigc_http_request, aigc_canvas_place, aigc_media_edit).
dsh-visual-plugin
by jyh20030112
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
dsh-short-video-studio
by fengyungithub
Adds image/video generation, canvas workspaces, ComfyUI workflow registration, media assets, and scene-production skills to DSH.
dsh-vision-proxy
by flyvhidbwo
DeepSeek brain + automatic image transcription — proxies attached images to a VLM (DashScope qwen by default) and feeds the transcribed text back to DeepSeek.
dsh-cad
by lau-mars
deepseek harness 2D and 3D CAD plugin
dsh-vision-opencode
by poiuyjie
DSH plugin: Auto-convert images to text for pure-text LLMs (DeepSeek etc.) via any vision model. No need to switch your main model.
dsh-ocr-plugin
by crazy222123
Local OCR provider (RapidOCR fast + DeepSeek-OCR-2/llama.cpp deep) registered as 'ocr' service on the llm-deepseek adapter seam; converts images to text blocks before API send.
vision-exp-tile
by nicholas023
Splits large images into coordinate-labelled tiles and recognizes them through DeepSeek vision workflows.
