1. Conduid
  2. Marketplace
  3. #evaluation
Tag

MCP servers tagged evaluation

19 live MCP servers tagged "evaluation", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with evaluation.

19
live servers
51
avg trust
144
most stars
1

Hegelion

Hmbown

Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)

★ 117updated 6 months agoMIT
80
2

Interviewer

microsoft

Catch MCP server issues before your agents do.

★ 144updated 8 months agoMIT
71
3

Kiln

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

64
4

Ragbits

Building blocks for rapid development of GenAI applications

64
5

Opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

64
6

Deepfabric

Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline

59
7

Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

59
8

As A Judge

OtherVibes

MCP as a Judge is a behavioral MCP that strengthens AI coding assistants by requiring explicit LLM evaluations

★ 16updated 8 months agoMIT
58
9

Primitive Bench

The marketplace for verifiable AI outcomes: state a task, get a fixed price upfront, and Primitive Bench routes across tools to complete it, refunded if the outcome isn't delivere…

52
10

Hypha

Harness-oriented agent system framework for production-grade LLM agent applications

52
11

Host

LeadFarmer-ai

MCP client manager: stdio + SSE transports, exp-backoff reconnect, server registry, built on @modelcontextprotocol/sdk

★ 12updated a year ago
46
12

mcpbr-cli

git+greynewell

Model Context Protocol Benchmark Runner - CLI tool for evaluating MCP servers

39
13

AI Testing MCP

MCP server for comprehensive AI testing, evaluation, and quality assurance

39
14

LLM Wiki Agent

Convert raw content into an interlinked Obsidian wiki using this Kotlin MCP server and Claude Code. Build a knowledge base without RAG indexing.

39
15

Mcplab MCP Server

MCP server that exposes MCPLab evaluation tools — query runs, results, and traces via the Model Context Protocol

37
16

Mcpbr CLI

Model Context Protocol Benchmark Runner - CLI tool for evaluating MCP servers

37
17

Dyno

Put your MCP server on the dyno — holistic, LLM-driven analysis of efficiency, cost, context-bloat, correctness, and reliability, with rigorous before/after error bars.

37
18

Oh My Field

Field-fit agents to real work. Turn tacit know-how into reusable capabilities.

34
19

Assay

Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.

34
← Previous Page 1 of 1 Next →