- Conduid
- Marketplace
- #benchmarking
MCP servers tagged benchmarking
25 live MCP servers tagged "benchmarking", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with benchmarking.
Hegelion
Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)
Goku
Goku is an HTTP load testing application written in Rust
Mcpbench
MCP server load tester: concurrent JSON-RPC, per-tool p50/p95/p99, stdio+HTTP/SSE transports, CI regression gating
Fastapi Bitnet
Running Microsoft's BitNet inference framework via FastAPI, Uvicorn and Docker.
LogLead
LogLead performs log loading, log enhancement, log feature engineering, log analysis, log anomaly detection also via MCP-server.
Context Diamond
Auditable context capsules for LLM handoffs, coding agents, and OpenCode MCP workflows.
Decibench
The open testing standard for voice AI agents. Deterministic + semantic + RAG augmented evaluation. Local first. Zero telemetry.
Blender Agent Studio
Codex plugin for reproducible Blender modeling, validation, animation, MCP tooling, and agent benchmarking
redline
Zero-infrastructure QA agent — k6 performance + Playwright functional testing with red/green baselines and a deterministic, model-free deploy gate. Clone and run.
Cpt AI
AI-powered benchmark regression analyzer with CLI, REST API, and interactive Dashboard. Compares cloud performance runs via OpenSearch + LLM, with automated regression detection.
Agent Smith 42
LLM code-agent framework — Thought→Code→Observation loop, process-isolated sandbox, MBPP + SWE-bench tracks, MCP tools, multi-model OpenRouter benchmarks. 42 School project.
Tne SDK
TNE-SDK - The Official Python SDK for The Null Epoch - an AI-only MMO where LLMs play autonomously. Connect any model (OpenAI, Anthropic, Ollama, vLLM, etc.) via MCP, TUI launcher…
Token Enhancer
Reduce web pages to clean text before they reach your AI agent’s context window, saving tokens and cutting noise
Claude Code For AI Engineers
Open-source preview of *Claude Code for AI Engineers* - a methodology-first skill pack for RAG eval, agent debugging, MCP servers, paper reproduction, and benchmark reporting. Ful…
Nf Runinsights
Nextflow plugin for cross-run benchmarking. Records per-process metrics for every run into a local or shared history store, compares runs against their own history, and flags regr…
Maverik
Benchmark, compare, and cost-predict your MCP agents — JMeter for AI agents.
Kaepsele
Full-stack AI Chatbot deployment for universities. Includes RAG, secure Code Interpreter, vLLM inference, and educational prompt libraries.
Tracebench
RCAgentBench: a benchmark measuring how tool design and agent instructions affect LLM agents doing root cause analysis on distributed traces.
Pigeon Evals
A End-To-End RAG Pipeline that includes Evaluations, iterations, and swappable components. At its core it allows users to be able to try different embedding models and techniques.
Linkedin Repo Role Recommender Agent
An AI agent built with MCP and LangGraph that analyzes GitHub repositories of LinkedIn users and recommends the top 3 most suitable IT roles for them.
Cortex Mem
Cortex — persistent AI memory for Claude Desktop & Claude Code. Local MCP server with a typed knowledge graph, contradiction detection & background scheduler. 20/20 benchmark, 27/…
Memvid
🧠 Experience advanced memory management with Memvid, a tool for interacting with your knowledge base using LLMs and smart recall features.
Halflist
CI-first conformance testing and benchmarking CLI for MCP servers. Lint your MCP server before your users do.
AI Agent Testing
A research project to measure AI agent robustness. Contains automated testing pipelines and a benchmarking methodology developed to audit Agentic AI architectures for complex reas…
Helios LLM Orchestrator
Local-first MCP multi-LLM orchestrator with task decomposition and benchmark-guided model routing through OpenRouter.