- Conduid
- Marketplace
- #ai-evaluation
MCP servers tagged ai-evaluation
17 live MCP servers tagged "ai-evaluation", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with ai-evaluation.
Headroom
More human oversight can make an AI agent less safe. Headroom is a human-in-the-loop firewall for coding agents that measures when to trust the human: oversight as resource alloca…
Nasde Toolkit
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
Grandjury
Pluralistic human evaluation infrastructure for AI in production
Mdb API Layer
A pytest-based API testing framework for the TMDB REST API, showcasing modern SDET: end-to-end integration tests, Pact contract validation, Dockerized parallel runs, Kubernetes-ba…
Goodbotbadbot
The MCP server behind goodbotbad.bot, where the crowd rules AI transcripts good bot or bad bot.
Harneloop
Build self-evolving AI agent harnesses with portable harness units, artifact-aware testing, trace-backed diagnosis, and evidence-gated promotion.
Agentic AI Learning Roadmap
A practical 12-week Agentic AI learning roadmap covering LLM foundations, tool calling, RAG, memory, MCP, multi-agent systems, evaluation, security, and production deployment.
Awesome Agentic AI
Curated Agentic AI resources: AI agent frameworks, tutorials, courses, projects, MCP tools, research papers, evaluation, security, and production guides.
Agent Skill Forge
Spec-driven AI agent skill generator and verification pipeline for building, testing, repairing, and packaging production-ready agent skills.
Multivon MCP
MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.
Intent Eval Lab
Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilot CLI, Codex CLI).
J Rig Skill Binary Eval
Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, tri…
Leadline
Evidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the ho…
Meva Health AI
Open-source framework for evaluating whether AI agents retrieve and faithfully use medical evidence from synthetic FHIR records.
AI Engineering Learning
An eight-day, agent-first AI engineering curriculum for non-technical learners.
Moses
MO§ES™ — Enterprise AI Operator Evaluations. Marketing site + MCP server. mos2es.org
Learn Agent Harness
No description published.