- Conduid
- Marketplace
- #benchmark
MCP servers tagged benchmark
24 live MCP servers tagged "benchmark", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with benchmark.
AI Infra Guard
A full-stack AI Red Teaming platform securing AI ecosystems via AI Infra scan, MCP scan, Agent skills scan, and LLM jailbreak evaluation.
Hotpath Rs
Rust Performance Profiler & Channels Monitoring Toolkit (TUI, MCP)
Aenvironment
Standardized environment infrastructure for Agentic AI development.
Goku
Goku is an HTTP load testing application written in Rust
Mcpmark
MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.
Mcpbench
The evaluation benchmark on MCP servers
Mcptoolbenchpp
MCPToolBench++ MCP Model Context Protocol Tool Use Benchmark on AI Agent and Model Tool Use Ability
Primitive Bench
The marketplace for verifiable AI outcomes: state a task, get a fixed price upfront, and Primitive Bench routes across tools to complete it, refunded if the outcome isn't delivere…
Spmind
ICML 2026 autonomous AI agent for end-to-end spatial proteomics analysis, with SP-Bench for agentic multiplexed-imaging workflows.
mcpbr-cli
Model Context Protocol Benchmark Runner - CLI tool for evaluating MCP servers
MCPSecBench
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
Civ6 MCP
An MCP server that lets LLM agents play Civilization VI.
Awesome AI Gateway
⚡ Awesome AI Gateway — curated comparison of 100+ AI gateways & LLM proxies (LiteLLM, OpenRouter, Portkey, Kong, Higress, new-api, Bifrost) by cost, security, compliance & self-ho…
ADR
ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
Stress
Stress testing tool for MCP servers
Mcpbr CLI
Model Context Protocol Benchmark Runner - CLI tool for evaluating MCP servers
Larkx
AI codebase indexer and MCP server for Claude Code, Cursor, and Copilot. Pre-index your project into a compact graph and measure real token savings with `larkx bench`.
Vitals
Vital signs for your MCP server: inspect capabilities, benchmark tool-call latency (p50/p95/p99), and assert health in CI. The ab/k6/pytest for MCP servers — non-interactive, scri…
Hlido MCP Guard
Check an MCP server's independent Hlido trust verdict before it starts — a zero-dependency, no-API-key runtime guardrail for Claude Code, Cursor, Cline, and any stdio MCP client.
Mcpscope
A local-first workbench for developing, inspecting, and benchmarking MCP servers against local (LM Studio, Ollama) or remote (OpenRouter) models — Web UI, CLI, and MCP interface o…
Verigent MCP Server
MCP server for Verigent — AI agent verification & counterparty due diligence. Check who you're transacting with, carry your own VG credential, flag bad actors.
Coder Bench
Benchmark tool for measuring MCP server effectiveness in LLM-assisted development
Meatloaf
🥩 AI Agent Sandbox Runtime — Build. Run. Test. Ship. Give AI agents a body to interact with the digital world.
ContextTax
Measure the context-window tax an MCP server charges your AI agent — the real token cost of its tool schemas + responses. Ground truth (Anthropic count_tokens) or keyless estimate…