1. Conduid
  2. Marketplace
  3. #inference
Tag

MCP servers tagged inference

13 live MCP servers tagged "inference", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with inference.

13
live servers
48
avg trust
458
most stars
1

Vllm Mlx

waybarrios

OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal su…

★ 458updated 6 months ago
69
2

client

git+huggingface

Use Hugging Face with JavaScript

64
3

Fastapi Bitnet

grctest

Running Microsoft's BitNet inference framework via FastAPI, Uvicorn and Docker.

★ 36updated a year agoMIT
60
4

Mlxstudio

MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)

59
5

Huggingface MCP Server

shreyaskarnik

MCP Server for HuggingFace inference endpoints with custom LoRA and story generation

★ 68updated a year agoMIT
55
6

Flama

The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native…

52
7

Quickstart Streaming Agents

Build, deploy, and orchestrate event-driven agents natively on Apache Flink® and Apache Kafka®

44
8

Rasputin Memory

The ultimate memory backend for OpenClaw and Claude Code. Persistent conversation memory with LLM fact extraction, foundation-model reranking, and 77.7% LoCoMo accuracy. MCP nativ…

39
9

Local AI

Unified MCP server for managing local model runtimes (Ollama, LM Studio, and more): provider-agnostic discovery, lifecycle, hardware-fit, and delegated inference.

37
10

Capix

Capix MCP Server — deploy and manage private LLM instances, compute, websites, and verified workloads on the Capix network from any AI coding agent that supports MCP.

37
11

TARILIO

AI-powered Information Retrieval with integrated AI Assistant, MCP Client, Local LLM server

34
12

VrtxOmega/Ollama-Omega

Official Ollama MCP Server. Exposes ollama_chat, ollama_generate, ollama_pull_model, ollama_list_models and ollama_show_model tools for advanced AI interactions.

34
13

Diffusion LLM MCP

FastMCP fleet MCP server for diffusion LMs (dLLM). DiffusionGemma on Goliath RTX 4090 — batch inference, HLE-shaped reasoning, ~200–400 tok/s. Doc phase; llama-diffusion-cli sidec…

34
← Previous Page 1 of 1 Next →