Make DeepSeek Harness see images: multimodal models vs the vision plugin

Paste a screenshot into stock DeepSeek and you only get an error — its models are text-only. Two real fixes, rebuilt frame by frame from a Bilibili course: connect Z.ai's GLM vision model, or install the Vision Router plugin that ships free vision models.

Last updated: 2026-09-17

Paste an image into a fresh DeepSeek Harness chat and the reply comes back fast — and negative: the current model does not support images. That is not a settings bug. DeepSeek's own models are text models that accept written input only, so screenshots, photos and diagrams all hit the same wall.

This guide rebuilds the vision-and-multimodal episode of 博学谷's Bilibili course frame by frame. The presenter demos two independent fixes — connecting a multimodal model from another provider, and installing a vision plugin — and we kept a screenshot of every step, each linking back to the exact second of the source video.

The 30-second version

Two ways to make DSH process images. Way 1: swap the model — add a provider under Settings → Models; the demo picks Z.ai (Zhipu GLM), a fresh account's free resource pack covers it, and the image goes to a vision model like GLM-5V-Turbo. Way 2: install a plugin — dsh-vision-router from the plugin market ships free vision models, so you keep your DeepSeek text model and the plugin does the seeing.

Both routes end the same way: the header screenshot that was rejected before comes back fully described — DeepSeek's fish-shaped logo on the left, the 探索未至之境 wordmark in the middle, a preview tag on the right. A plugin is not magic: behind the scenes it also calls a multimodal model.

Tutorial slide contrasting the two ways DeepSeek Harness can process visual material: a multimodal model that sees images itself on the left, and an installed vision or multimodal plugin on the right, both meeting at the DSH core
The course's own map of this episode: bring a model that sees, or install a plugin that sees for it.Watch at 4:20

Give DSH eyes, two ways, seven steps

  1. 1

    Paste an image into stock DSH and watch it bounce

    The demo screenshots DSH's own header — logo left, 探索未至之境 wordmark center — pastes it into the composer, and asks what's in the picture with DeepSeek-V4-Pro selected. The reply is an error toast: the current model does not support images. DeepSeek's two built-in models are text models that accept written input only, so switching between them changes nothing.

    DeepSeek Harness composer with the app header screenshot pasted and the question what's-in-this-image typed in Chinese, while the text model DeepSeek-V4-Pro is selected
    First try: the header screenshot goes in with a plain what-do-you-see question.Watch at 0:23
    Error toast in DeepSeek Harness saying the current model does not support images and to pick an image-capable model — the built-in text model rejecting a pasted screenshot
    Stock DeepSeek models only take text — the image is never even processed.Watch at 0:33
  2. 2

    Open Settings → Models and add a provider that sees

    From the gear at the bottom left, open Settings → Models. Next to the built-in DeepSeek entry sits an add-provider catalog: amazon-bedrock, openai, google, minimax and many more. The episode picks z-ai — Zhipu's Z.ai, home of the GLM models — which wants either its TokenPlan package or a fresh account's free credits.

    Provider dropdown in DeepSeek Harness model settings listing providers from amazon-bedrock to z-ai, with the Z.ai entry highlighted for its GLM vision models
    The catalog is long; this episode scrolls to the bottom and picks z-ai (Zhipu GLM).Watch at 1:41
  3. 3

    Grab a free API key and paste it into DSH

    The course notes link straight to BigModel, Zhipu's console: registering with a new phone number grants free resource packs (check Finance → Resource packs — enough to experiment). On the API Key page, create a key — any name works — then copy it into DSH's ZAI provider form and hit save. Z.ai's models unlock immediately.

    BigModel API key page with the create-key dialog open, where any nickname works and confirming issues the key that DeepSeek Harness will store
    A new phone number gets free resource packs; the API key itself is one dialog away.Watch at 3:01
    ZAI provider form inside DeepSeek Harness settings with the API key pasted as a masked value, cursor resting on the save button
    Paste, save, done — that is the whole handshake.Watch at 3:13
  4. 4

    Switch to a GLM vision model and send the image again

    Reopen the model picker: a z-ai group now sits under the stock DeepSeek entries. Vision belongs to single models, not to the provider — the demo first tries GLM-5.2 and gets the same not-supported toast, then moves to GLM-5V-Turbo, which accepts the upload. The answer reads the screenshot element by element: black fish-shaped logo left, 探索未至之境 text center, preview tag right.

    DeepSeek Harness model picker showing a new z-ai group with GLM-4.5-Air, GLM-4.7, GLM-5-Turbo, GLM-5.1 and GLM-5.2 beneath the built-in DeepSeek models
    The provider is in — now pick carefully, because not every GLM model sees.Watch at 3:23
    GLM-5V-Turbo answering inside DeepSeek Harness: the pasted header image holds a black fish-shaped DeepSeek logo on the left, the Chinese wordmark in the middle and a preview tag on the right
    Way 1 works: the vision model narrates logo, wordmark and tag in one pass.Watch at 4:00
  5. 5

    Install the Vision Router plugin from the market

    Way 2 keeps your text model. Open Settings → Plugin market and pick the vision-and-multimodal category. The episode installs dsh-vision-router: it gives DSH image understanding and — the decisive part — bundles free vision models, so no extra API key on any platform is needed. Click install, then confirm.

    Install confirmation dialog for the dsh-vision-router plugin in the DeepSeek Harness plugin market — the image-recognition plugin with built-in free vision models
    dsh-vision-router is the pick: free built-in vision models, zero extra keys.Watch at 5:13
  6. 6

    Restart DSH, then check the plugin's vision settings

    The installer asks for a restart. The presenter skips the restart button — it has misbehaved on his setup before — and instead stops the process with Ctrl+C in the terminal, re-runs the start command, and refreshes the browser. Settings → Plugins now lists Vision Router (image recognition), where you can assign the vision model it should call; built-in free fallback models and a usage limit keep casual use free.

    Vision Router plugin settings in DeepSeek Harness showing the vision-model provider picker and the built-in free vision models that fall back when nothing is configured
    After the restart the plugin shows up in settings — free fallback models included, no key required.Watch at 6:10
  7. 7

    Pick the vision-tagged model variant and ask again

    Still on text-only DeepSeek-V4-Pro, open the model picker: the plugin added a second DeepSeek group whose entries carry a vision badge — same models, plus automatic image recognition. Choose the variant, paste the same header screenshot, and ask what's in the picture again. This time the answer dissects the image into three blocks: fish-shaped logo left, wordmark center, preview tag right. DSH can see.

    DeepSeek Harness model picker with a plugin-added second DeepSeek group whose vision-badged DeepSeek-V4-Flash and DeepSeek-V4-Pro entries sit next to the original text-only ones
    Same names, new sense: the plugin's variants bolt image recognition onto DeepSeek-V4-Pro.Watch at 6:45
    DeepSeek Harness answering the what's-in-this-image question through the Vision Router plugin, splitting the screenshot into three horizontal blocks: fish logo left, Chinese wordmark center, preview tag right
    Way 2 works too — the text model stayed, the plugin's vision model did the reading.Watch at 7:18

Image recognition in DSH: quick answers

Models, keys, costs and restarts — the questions this episode actually answers.

Does DeepSeek Harness support vision models?

Not out of the box — the built-in DeepSeek models are text-only, as the error toast in step one shows. You add vision in one of two ways: connect a provider whose models are multimodal (the demo uses Z.ai and sends to GLM-5V-Turbo), or install the Vision Router plugin, which routes images to vision models — including free built-in ones — while you keep chatting with your text model.

Do I have to pay for the GLM key or the vision plugin?

No. Registering BigModel, Zhipu's console, with a fresh phone number grants free resource packs — the demo checks Finance → Resource packs and finds enough to experiment. The Vision Router plugin itself ships free vision models, so its fallback path costs nothing either. Heavy use is where the TokenPlan package or your own key comes in.

Why did GLM-5.2 still refuse my image when GLM-5V-Turbo worked?

Because vision belongs to individual models, not to the provider. Z.ai's catalog mixes text models (GLM-4.7, GLM-5.1, GLM-5.2…) with the GLM-5V vision family. Pick a text model and you get the same image-not-supported toast; pick GLM-5V-Turbo and the composer accepts the image immediately.

The plugin asks to restart — should I click its Restart button?

You can, but the presenter avoids it: restarts have glitched on his setup before. His route: Ctrl+C in the terminal to stop dsh, run the start command again, refresh the browser — the plugin then appears under Settings → Plugins. If you do use the button and something breaks, ask DSH itself to fix it.

Related DeepSeek Harness guides

From picking the vision plugin to swapping models — the adjacent reads.

Sources & credits

Frames in this guide are captured from the Bilibili course below and remain the property of its creator — each image links back to the exact second it was taken.

DSH Plugins is an independent community directory of DeepSeek Harness plugins. Not affiliated with or endorsed by DeepSeek. Third-party plugins are not security-audited — review the source before installing.

New DeepSeek Harness plugins, weekly. No spam.