1. Conduid
  2. Marketplace
  3. #benchmarking
Tag

MCP servers tagged benchmarking

25 live MCP servers tagged "benchmarking", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with benchmarking.

25
live servers
41
avg trust
241
most stars
1

Hegelion

Hmbown

Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)

★ 117updated 6 months agoMIT
80
2

Goku

jcaromiq

Goku is an HTTP load testing application written in Rust

★ 145updated 9 months agoMIT
77
3

Mcpbench

modelscope

MCP server load tester: concurrent JSON-RPC, per-tool p50/p95/p99, stdio+HTTP/SSE transports, CI regression gating

★ 241updated a year agoApache-2.0
74
4

Fastapi Bitnet

grctest

Running Microsoft's BitNet inference framework via FastAPI, Uvicorn and Docker.

★ 36updated a year agoMIT
60
5

LogLead

LogLead performs log loading, log enhancement, log feature engineering, log analysis, log anomaly detection also via MCP-server.

39
6

Context Diamond

Auditable context capsules for LLM handoffs, coding agents, and OpenCode MCP workflows.

39
7

Decibench

The open testing standard for voice AI agents. Deterministic + semantic + RAG augmented evaluation. Local first. Zero telemetry.

39
8

Blender Agent Studio

Codex plugin for reproducible Blender modeling, validation, animation, MCP tooling, and agent benchmarking

39
9

redline

Zero-infrastructure QA agent — k6 performance + Playwright functional testing with red/green baselines and a deterministic, model-free deploy gate. Clone and run.

★ 2
37
10

Cpt AI

sayalibhavsar

AI-powered benchmark regression analyzer with CLI, REST API, and interactive Dashboard. Compares cloud performance runs via OpenSearch + LLM, with automated regression detection.

★ 1
34
11

Agent Smith 42

Sabofxx

LLM code-agent framework — Thought→Code→Observation loop, process-isolated sandbox, MBPP + SWE-bench tracks, MCP tools, multi-model OpenRouter benchmarks. 42 School project.

★ 1
34
12

Tne SDK

TNE-SDK - The Official Python SDK for The Null Epoch - an AI-only MMO where LLMs play autonomously. Connect any model (OpenAI, Anthropic, Ollama, vLLM, etc.) via MCP, TUI launcher…

34
13

Token Enhancer

Reduce web pages to clean text before they reach your AI agent’s context window, saving tokens and cutting noise

34
14

Claude Code For AI Engineers

Open-source preview of *Claude Code for AI Engineers* - a methodology-first skill pack for RAG eval, agent debugging, MCP servers, paper reproduction, and benchmark reporting. Ful…

34
15

Nf Runinsights

Nextflow plugin for cross-run benchmarking. Records per-process metrics for every run into a local or shared history store, compares runs against their own history, and flags regr…

34
16

Maverik

Benchmark, compare, and cost-predict your MCP agents — JMeter for AI agents.

34
17

Kaepsele

Full-stack AI Chatbot deployment for universities. Includes RAG, secure Code Interpreter, vLLM inference, and educational prompt libraries.

34
18

Tracebench

RCAgentBench: a benchmark measuring how tool design and agent instructions affect LLM agents doing root cause analysis on distributed traces.

34
19

Pigeon Evals

A End-To-End RAG Pipeline that includes Evaluations, iterations, and swappable components. At its core it allows users to be able to try different embedding models and techniques.

34
20

Linkedin Repo Role Recommender Agent

An AI agent built with MCP and LangGraph that analyzes GitHub repositories of LinkedIn users and recommends the top 3 most suitable IT roles for them.

34
21

Cortex Mem

Cortex — persistent AI memory for Claude Desktop & Claude Code. Local MCP server with a typed knowledge graph, contradiction detection & background scheduler. 20/20 benchmark, 27/…

34
22

Memvid

🧠 Experience advanced memory management with Memvid, a tool for interacting with your knowledge base using LLMs and smart recall features.

34
23

Halflist

CI-first conformance testing and benchmarking CLI for MCP servers. Lint your MCP server before your users do.

34
24

AI Agent Testing

A research project to measure AI agent robustness. Contains automated testing pipelines and a benchmarking methodology developed to audit Agentic AI architectures for complex reas…

34
25

Helios LLM Orchestrator

Erfouni

Local-first MCP multi-LLM orchestrator with task decomposition and benchmark-guided model routing through OpenRouter.

34
← Previous Page 1 of 1 Next →