- Conduid
- Marketplace
- #llm-as-judge
MCP servers tagged llm-as-judge
11 live MCP servers tagged "llm-as-judge", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with llm-as-judge.
Tracely AI
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
Tracely
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
AI Engineer Notebooks
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, ag…
Agentevals
agentevals is a framework-agnostic evaluations solution based on OpenTelemetry traces
BestTester
Production-grade Playwright + TypeScript QA framework with AI-powered testing, LLM-as-Judge evaluation, MCP server, 7 CLI agents, security fuzzing, CI/CD pipelines, Jira sync, and…
Classifier Evals MCP Server
MCP server exposing classifier evaluation tools
Prism Verify
prism-verify — runtime LLM verifier (family-different, reasoning-stripped, multi-lens, signed Ed25519 receipts). Zero-prerequisite npx install via a verified binary launcher.
Omegaprompt
The overfit gate for your prompts: re-test the winning prompt on held-out examples it never tuned on, and block the ship if it doesn't generalize. Sits on top of promptfoo/DSPy. C…
Clawdmarket
Autonomous agent-to-agent marketplace with live Karpathy loop self-improvement. Agents discover, hire, benchmark, and evolve programmatically. MPP/x402/MCP. No humans in the loop.
Humanloopbench
Human-in-the-loop LLM Failure Detection & Benchmark Toolkit — Auto-flag LLM failures (hallucination, refusal, factual error) with LLM-as-Judge explanations. Confirm them in a revi…
Ferryman MCP
A local-first MCP host with pluggable skills, multi-provider routing, and multi-channel I/O. Kotlin. Plus a Python eval harness with rule + LLM-judge scorers over a human-authored…