- Conduid
- Marketplace
- #agent-evaluation
MCP servers tagged agent-evaluation
6 live MCP servers tagged "agent-evaluation", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with agent-evaluation.
Any Agent
A single interface to use and evaluate different agent frameworks
Eval View
Proof your AI agent still works. Regression testing with golden baselines, tool-call diffing, and output drift detection. MCP server + Claude Code skills. LangGraph, CrewAI, Anthr…
Trulens
Evaluation and Tracking for LLM Experiments and AI Agents
DM Code Agent
Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite eval and trace replay.
Coder Eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Hlido MCP Guard
Check an MCP server's independent Hlido trust verdict before it starts — a zero-dependency, no-API-key runtime guardrail for Claude Code, Cursor, Cline, and any stdio MCP client.