1. Conduid
  2. Marketplace
  3. #benchmark
Tag

MCP servers tagged benchmark

24 live MCP servers tagged "benchmark", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with benchmark.

24
live servers
50
avg trust
3.0K
most stars
1

AI Infra Guard

Tencent

A full-stack AI Red Teaming platform securing AI ecosystems via AI Infra scan, MCP scan, Agent skills scan, and LLM jailbreak evaluation.

★ 3.0Kupdated 6 months agoNOASSERTION
92
2

Hotpath Rs

pawurb

Rust Performance Profiler & Channels Monitoring Toolkit (TUI, MCP)

★ 1.3Kupdated 6 months agoMIT
92
3

Aenvironment

inclusionAI

Standardized environment infrastructure for Agentic AI development.

★ 254updated 6 months agoApache-2.0
80
4

Goku

jcaromiq

Goku is an HTTP load testing application written in Rust

★ 145updated 9 months agoMIT
77
5

Mcpmark

eval-sys

MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.

★ 388updated 7 months agoApache-2.0
74
6

Mcpbench

modelscope

The evaluation benchmark on MCP servers

★ 241updated 12 months agoApache-2.0
74
7

Mcptoolbenchpp

mcp-tool-bench

MCPToolBench++ MCP Model Context Protocol Tool Use Benchmark on AI Agent and Model Tool Use Ability

★ 41updated 8 months ago
53
8

Primitive Bench

The marketplace for verifiable AI outcomes: state a task, get a fixed price upfront, and Primitive Bench routes across tools to complete it, refunded if the outcome isn't delivere…

52
9

Spmind

ICML 2026 autonomous AI agent for end-to-end spatial proteomics analysis, with SP-Bench for agentic multiplexed-imaging workflows.

52
10

mcpbr-cli

git+greynewell

Model Context Protocol Benchmark Runner - CLI tool for evaluating MCP servers

39
11

MCPSecBench

MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols

39
12

Civ6 MCP

An MCP server that lets LLM agents play Civilization VI.

39
13

Awesome AI Gateway

⚡ Awesome AI Gateway — curated comparison of 100+ AI gateways & LLM proxies (LiteLLM, OpenRouter, Portkey, Kong, Higress, new-api, Bifrost) by cost, security, compliance & self-ho…

39
14

ADR

ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.

39
15

Stress

Stress testing tool for MCP servers

37
16

Mcpbr CLI

Model Context Protocol Benchmark Runner - CLI tool for evaluating MCP servers

37
17

Larkx

AI codebase indexer and MCP server for Claude Code, Cursor, and Copilot. Pre-index your project into a compact graph and measure real token savings with `larkx bench`.

37
18

Vitals

Vital signs for your MCP server: inspect capabilities, benchmark tool-call latency (p50/p95/p99), and assert health in CI. The ab/k6/pytest for MCP servers — non-interactive, scri…

37
19

Hlido MCP Guard

Check an MCP server's independent Hlido trust verdict before it starts — a zero-dependency, no-API-key runtime guardrail for Claude Code, Cursor, Cline, and any stdio MCP client.

37
20

Mcpscope

A local-first workbench for developing, inspecting, and benchmarking MCP servers against local (LM Studio, Ollama) or remote (OpenRouter) models — Web UI, CLI, and MCP interface o…

37
21

Verigent MCP Server

MCP server for Verigent — AI agent verification & counterparty due diligence. Check who you're transacting with, carry your own VG credential, flag bad actors.

37
22

Coder Bench

Benchmark tool for measuring MCP server effectiveness in LLM-assisted development

34
23

Meatloaf

🥩 AI Agent Sandbox Runtime — Build. Run. Test. Ship. Give AI agents a body to interact with the digital world.

34
24

ContextTax

Measure the context-window tax an MCP server charges your AI agent — the real token cost of its tool schemas + responses. Ground truth (Anthropic count_tokens) or keyless estimate…

34
← Previous Page 1 of 1 Next →