1. Conduid
  2. Marketplace
  3. #vision-language-model
Tag

MCP servers tagged vision-language-model

14 live MCP servers tagged "vision-language-model", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with vision-language-model.

14
live servers
37
avg trust
4
most stars
1

OGAM

The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and loc…

64
2

Images

MCP 图片分析服务器 — 在 AI 工具中直接分析本地图片、剪贴板截图或 Base64 图片数据

★ 1
39
3

Jy Crpg Bench

A long-horizon CRPG benchmark for frontier agents. Raw 320x200 frames, Traditional Chinese, isometric navigation, twelve books to find.

39
4

Visbridge

kasumaputu6633

Token-efficient vision capability layer for MCP clients (describe / OCR / inspect)

★ 4
37
5

Vlm MCP Server

MCP Server for VLM - A Model Context Protocol server providing vision/video analysis, configurable with any model provider (Chat Completions / Responses / Anthropic).

37
6

Blink Skill

Proactive screen awareness + Claude Vision assistance

34
7

Shadow Vision

Open-source MCP vision server that gives text-only LLMs and AI agents image understanding, OCR, visual analysis, UI inspection, and multimodal capabilities.

34
8

Her Eyes

带上她的眼睛 · Give a text-only LLM eyes — a single-binary MCP tool that lets agents like Claude Code / Codex call a vision model to extract structured key information from images.

34
9

Nvidia MCP

MCP server that gives Claude, Cursor and any AI agent free access to 100+ NVIDIA-hosted models (Nemotron, Llama, GPT-OSS, DeepSeek, vision, embeddings) with automatic task routing…

34
10

Claude Holo Unreal

Unreal Engine copilot for Claude Code — CLI + MCP server + skills + Python library. Drives the running UE editor via PythonScriptPlugin, falls back to H Company Holo3 vision groun…

34
11

Screen Use

browser-use, but for the entire desktop — give any AI Agent eyes and hands on Windows. MCP + SDK, VLM-optional, with visual loop, introspection and meta-learning.

34
12

Doc7

Convert any document—PDFs, Office files, scans, charts—into AI-ready Markdown for seamless search, quoting, and reasoning.

34
13

Adatile MCP

AdaTile-MCP: high-resolution image adaptive tiling MCP server for DeepSeek vision model (deepseek-v4-flash-vision-exp). L1-L6 pipeline (fastpath, saliency, adaptive tiling, Files…

34
14

Oai Chat

Multi-modal Chatbot based on OpenAI

34
← Previous Page 1 of 1 Next →