Vision plugins
DeepSeek Harness ships text-only by default. Vision plugins close that gap: they turn pasted screenshots into structured OCR evidence, answer questions about images, reconstruct UI layouts from pixels, and diff two screenshots so the agent can see what changed. If your workflow involves pasting error screens or design mockups into chat, start here.
42 plugins
1–30 of 42
Vision plugin FAQ
- How do DSH vision plugins work?
- Most register a view_image-style tool or intercept pasted images, run OCR and layout detection, then feed the model structured text evidence instead of raw pixels. The model itself stays text-only.
- Do vision plugins need an external API key?
- Some use local OCR engines and need nothing else; others call a hosted vision model. The install note on each plugin page says which it is.
- Which vision plugin should I install first?
- ModLens is the most-starred starting point for paste-an-image OCR. DSH Vision Toolkit covers more ground: image Q&A, long screenshots, and UI reconstruction.




















