- Conduid
- Marketplace
- #agent-benchmark
MCP servers tagged agent-benchmark
7 live MCP servers tagged "agent-benchmark", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with agent-benchmark.
Eval View
Proof your AI agent still works. Regression testing with golden baselines, tool-call diffing, and output drift detection. MCP server + Claude Code skills. LangGraph, CrewAI, Anthr…
Nasde Toolkit
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
Jy Crpg Bench
A long-horizon CRPG benchmark for frontier agents. Raw 320x200 frames, Traditional Chinese, isometric navigation, twelve books to find.
haoyifan/Silicon-Pantheon
host option.
Unreal Agent Benchmark
Evidence-based Unreal Engine 5.8 benchmark for complete AI game-building systems using Epic Unreal MCP, playable tasks, packaging, and launch proof.
Aobench
Role-aware, permission-enforced benchmark for LLM agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. A policy violation hard-fails the task. 88 tasks (10 ca…
Craftonomous
Craftonomous is an agent-agnostic Minecraft embodiment and evaluation substrate: an MCP-native body with a declared perception budget and reliability-tracked skills.