- Conduid
- Marketplace
- #llm-inference
MCP servers tagged llm-inference
13 live MCP servers tagged "llm-inference", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with llm-inference.
Lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Monocle
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written in Python.
Qurio
Qurio is a high-velocity AI knowledge workspace built for teams that demand more than basic chat. It supports generic providers. Highlights include Deep Research for complex tasks…
RuVector
RuVector is a High Performance, Real-Time, Self-Learning, Vector GNN, Memory DB built in Rust.
Optillm
Optimizing inference proxy for LLMs
Spiceai
A portable accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents.
Mindbridge MCP
MindBridge is an AI orchestration MCP server that lets any app talk to any LLM — OpenAI, Anthropic, DeepSeek, Ollama, and more — through a single unified API. Route queries, compa…
Sample Genai On Eks Starter Kit
A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: 🚀 AI Gateway (LiteLLM) 🤖 LLM Serving (v…
Atomic Chat
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.
Deeppowers
DEEPPOWERS is a Fully Homomorphic Encryption (FHE) framework built for MCP (Model Context Protocol), aiming to provide end-to-end privacy protection and high-efficiency computatio…
Captain Claw
Self-hosted framework for orchestrating fleets of specialist AI agents — ensemble reasoning and a full agentic coding pipeline, model-agnostic and local-friendly.
Tokio Prompt Orchestrator
Multi-core, Tokio-native orchestration for LLM pipelines.
Rai
CPU-only LLM inference engine in pure Rust — 4-bit quantized models, hand-written AVX2 kernels, speculative decoding, and a local HTTP/MCP server. No GPU, no Python runtime.