1. Conduid
  2. Marketplace
  3. #llm-inference
Tag

MCP servers tagged llm-inference

13 live MCP servers tagged "llm-inference", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with llm-inference.

13
live servers
61
avg trust
2.3K
most stars
1

Lemonade

lemonade-sdk

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

★ 2.3Kupdated 6 months agoApache-2.0
97
2

Monocle

monocle2ai

Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written in Python.

★ 70updated 6 months agoApache-2.0
72
3

Qurio

havingautism

Qurio is a high-velocity AI knowledge workspace built for teams that demand more than basic chat. It supports generic providers. Highlights include Deep Research for complex tasks…

★ 41updated 6 months agoNOASSERTION
67
4

RuVector

RuVector is a High Performance, Real-Time, Self-Learning, Vector GNN, Memory DB built in Rust.

64
5

Optillm

Optimizing inference proxy for LLMs

64
6

Spiceai

A portable accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents.

64
7

Mindbridge MCP

pinkpixel-dev

MindBridge is an AI orchestration MCP server that lets any app talk to any LLM — OpenAI, Anthropic, DeepSeek, Ollama, and more — through a single unified API. Route queries, compa…

★ 27updated a year agoMIT
62
8

Sample Genai On Eks Starter Kit

aws-samples

A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: 🚀 AI Gateway (LiteLLM) 🤖 LLM Serving (v…

★ 39updated 6 months agoMIT-0
61
9

Atomic Chat

Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.

59
10

Deeppowers

DEEPPOWERS is a Fully Homomorphic Encryption (FHE) framework built for MCP (Model Context Protocol), aiming to provide end-to-end privacy protection and high-efficiency computatio…

52
11

Captain Claw

Self-hosted framework for orchestrating fleets of specialist AI agents — ensemble reasoning and a full agentic coding pipeline, model-agnostic and local-friendly.

52
12

Tokio Prompt Orchestrator

Multi-core, Tokio-native orchestration for LLM pipelines.

44
13

Rai

CPU-only LLM inference engine in pure Rust — 4-bit quantized models, hand-written AVX2 kernels, speculative decoding, and a local HTTP/MCP server. No GPU, no Python runtime.

34
← Previous Page 1 of 1 Next →