- Conduid
- Marketplace
- #inference
MCP servers tagged inference
13 live MCP servers tagged "inference", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with inference.
Vllm Mlx
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal su…
client
Use Hugging Face with JavaScript
Fastapi Bitnet
Running Microsoft's BitNet inference framework via FastAPI, Uvicorn and Docker.
Mlxstudio
MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)
Huggingface MCP Server
MCP Server for HuggingFace inference endpoints with custom LoRA and story generation
Flama
The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native…
Quickstart Streaming Agents
Build, deploy, and orchestrate event-driven agents natively on Apache Flink® and Apache Kafka®
Rasputin Memory
The ultimate memory backend for OpenClaw and Claude Code. Persistent conversation memory with LLM fact extraction, foundation-model reranking, and 77.7% LoCoMo accuracy. MCP nativ…
Local AI
Unified MCP server for managing local model runtimes (Ollama, LM Studio, and more): provider-agnostic discovery, lifecycle, hardware-fit, and delegated inference.
Capix
Capix MCP Server — deploy and manage private LLM instances, compute, websites, and verified workloads on the Capix network from any AI coding agent that supports MCP.
TARILIO
AI-powered Information Retrieval with integrated AI Assistant, MCP Client, Local LLM server
VrtxOmega/Ollama-Omega
Official Ollama MCP Server. Exposes ollama_chat, ollama_generate, ollama_pull_model, ollama_list_models and ollama_show_model tools for advanced AI interactions.
Diffusion LLM MCP
FastMCP fleet MCP server for diffusion LMs (dLLM). DiffusionGemma on Goliath RTX 4090 — batch inference, HLE-shaped reasoning, ~200–400 tok/s. Doc phase; llama-diffusion-cli sidec…