- Conduid
- Marketplace
- #vision
MCP servers tagged vision
25 live MCP servers tagged "vision", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with vision.
mediar-ai/screenpipe
run agents that work for you in the background based on what you do
Luma MCP
Multi-Model Visual Understanding MCP Server, GLM-4.6V, DeepSeek-OCR (free), and Qwen3-VL-Flash. Provide visual processing capabilities for AI coding models that do not support ima…
jcdickinson/simplemem
SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
mediar-ai/screenpipe
aware AI agents through a NextJS plugin ecosystem.
Oscribe
If you can see it, Oscribe can click it.
Achatbot
An open source chat bot architecture for voice/vision (and multimodal) assistants, local(CPU/GPU bound) and remote(I/O bound) to run.
FlowVision
FlowVision,一键分析项目代码 · 自动生成架构流程图 · 可视化编辑 · MCP 实时同步
Deepseek Eyes
给 DeepSeek 装上眼睛 — MCP Server + 通义千问VL, 剪贴板图片→视觉模型→文字描述 / Give DeepSeek the ability to see images via clipboard + Qwen-VL
Opencode Minimax Easy Vision
OpenCode plugin that restores the paste-and-ask workflow for text-only models by saving pasted images and injecting MCP tool instructions
image-tiler-server
MCP server that splits large images into optimally-sized tiles for LLM vision (Claude, OpenAI, Gemini, Gemini 3)
Screenshot X402 CLI
CLI for screenshot-x402: MCP Streamable HTTP client with x402 USDC payments — health, screenshots, and vision analysis.
Pi Zai
Unofficial pi package that exposes Z.ai MCP server tools for web search, URL reading, repository reading, and vision workflows.
Vision Server
MCP stdio server for image recognition via an existing vision model
Doubao Vision MCP Server
MCP server for Doubao vision models via Volcengine Ark API
LLM Vision
A TypeScript MCP server that gives text-only LLMs image understanding through StepFun vision models.
Image Vision
MCP server for image recognition via vision models (extensible to multiple providers), with an HTTP image-intercepting proxy that lets text-only LLMs accept image input in Claude…
Vision Foundation
Vision understanding MCP service powered by local SmolVLM / SmolVLM2 / MiniCPM-V via llama.cpp
Openai Vision MCP Server
MCP server for secure, bounded OpenAI-compatible vision analysis
Omni Context
Universal vision & file processor MCP server for LLMs: OCR screenshots, read code/text files, and dump zip archives as a clean sequential file-tree dump.
Agentic Vision
Persistent visual memory for AI agents — capture screenshots, embed with CLIP ViT-B/32, compare, recall. MCP server + Rust core library.
404
TEAM 404 : Project For Prabal (Hackathon) FarmGenius is an advanced, multi-agent agricultural assistant designed for Indian farmers, agri-entrepreneurs, and researchers. It combin…
keiver/image-tiler-mcp-server
based tile classification.
Minicpm Vision MCP
为 DeepSeek V4.0 等单模态大模型装上眼睛和耳朵 —— 本地视觉+音频 MCP 服务器,基于 Ollama + MiniCPM-V 4.6 + faster-whisper,图片描述·视频分析·语音转文字
Agent Vision MCP
An MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model —…
DS VISION MCP SKILL
DS-VISION V3: auditable visual evidence for coding agents via MCP (visual evidence for agents without native vision).