1. Conduid
  2. Marketplace
  3. #ai-evaluation
Tag

MCP servers tagged ai-evaluation

17 live MCP servers tagged "ai-evaluation", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with ai-evaluation.

17
live servers
37
avg trust
69K
most stars
1

Headroom

chopratejas

More human oversight can make an AI agent less safe. Headroom is a human-in-the-loop firewall for coding agents that measures when to trust the human: oversight as resource alloca…

★ 69Kupdated 6 months agoApache-2.0
87
2

Nasde Toolkit

CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.

39
3

Grandjury

Pluralistic human evaluation infrastructure for AI in production

34
4

Mdb API Layer

A pytest-based API testing framework for the TMDB REST API, showcasing modern SDET: end-to-end integration tests, Pact contract validation, Dockerized parallel runs, Kubernetes-ba…

34
5

Goodbotbadbot

The MCP server behind goodbotbad.bot, where the crowd rules AI transcripts good bot or bad bot.

34
6

Harneloop

Build self-evolving AI agent harnesses with portable harness units, artifact-aware testing, trace-backed diagnosis, and evidence-gated promotion.

34
7

Agentic AI Learning Roadmap

A practical 12-week Agentic AI learning roadmap covering LLM foundations, tool calling, RAG, memory, MCP, multi-agent systems, evaluation, security, and production deployment.

34
8

Awesome Agentic AI

Curated Agentic AI resources: AI agent frameworks, tutorials, courses, projects, MCP tools, research papers, evaluation, security, and production guides.

34
9

Agent Skill Forge

Spec-driven AI agent skill generator and verification pipeline for building, testing, repairing, and packaging production-ready agent skills.

34
10

Multivon MCP

MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.

34
11

Intent Eval Lab

Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilot CLI, Codex CLI).

34
12

J Rig Skill Binary Eval

Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, tri…

34
13

Leadline

Evidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the ho…

34
14

Meva Health AI

Open-source framework for evaluating whether AI agents retrieve and faithfully use medical evidence from synthetic FHIR records.

34
15

AI Engineering Learning

An eight-day, agent-first AI engineering curriculum for non-technical learners.

34
16

Moses

SunrisesIllNeverSee

MO§ES™ — Enterprise AI Operator Evaluations. Marketing site + MCP server. mos2es.org

34
17

Learn Agent Harness

No description published.

30
← Previous Page 1 of 1 Next →