- Conduid
- Marketplace
- #evaluation
MCP servers tagged evaluation
19 live MCP servers tagged "evaluation", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with evaluation.
Hegelion
Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)
Interviewer
Catch MCP server issues before your agents do.
Kiln
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Ragbits
Building blocks for rapid development of GenAI applications
Opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Deepfabric
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
Tracely
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
As A Judge
MCP as a Judge is a behavioral MCP that strengthens AI coding assistants by requiring explicit LLM evaluations
Primitive Bench
The marketplace for verifiable AI outcomes: state a task, get a fixed price upfront, and Primitive Bench routes across tools to complete it, refunded if the outcome isn't delivere…
Hypha
Harness-oriented agent system framework for production-grade LLM agent applications
Host
MCP client manager: stdio + SSE transports, exp-backoff reconnect, server registry, built on @modelcontextprotocol/sdk
mcpbr-cli
Model Context Protocol Benchmark Runner - CLI tool for evaluating MCP servers
AI Testing MCP
MCP server for comprehensive AI testing, evaluation, and quality assurance
LLM Wiki Agent
Convert raw content into an interlinked Obsidian wiki using this Kotlin MCP server and Claude Code. Build a knowledge base without RAG indexing.
Mcplab MCP Server
MCP server that exposes MCPLab evaluation tools — query runs, results, and traces via the Model Context Protocol
Mcpbr CLI
Model Context Protocol Benchmark Runner - CLI tool for evaluating MCP servers
Dyno
Put your MCP server on the dyno — holistic, LLM-driven analysis of efficiency, cost, context-bloat, correctness, and reliability, with rigorous before/after error bars.
Oh My Field
Field-fit agents to real work. Turn tacit know-how into reusable capabilities.
Assay
Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.