1. Conduid
  2. Marketplace
  3. #llm-as-judge
Tag

MCP servers tagged llm-as-judge

11 live MCP servers tagged "llm-as-judge", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with llm-as-judge.

11
live servers
44
avg trust
17
most stars
1

Tracely AI

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

64
2

Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

59
3

AI Engineer Notebooks

Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set — model APIs, structured output, tool calling, RAG, evals-as-the-spine, ag…

59
4

Agentevals

agentevals is a framework-agnostic evaluations solution based on OpenTelemetry traces

52
5

BestTester

Production-grade Playwright + TypeScript QA framework with AI-powered testing, LLM-as-Judge evaluation, MCP server, 7 CLI agents, security fuzzing, CI/CD pipelines, Jira sync, and…

★ 17
39
6

Classifier Evals MCP Server

MCP server exposing classifier evaluation tools

37
7

Prism Verify

prism-verify — runtime LLM verifier (family-different, reasoning-stripped, multi-lens, signed Ed25519 receipts). Zero-prerequisite npx install via a verified binary launcher.

37
8

Omegaprompt

The overfit gate for your prompts: re-test the winning prompt on held-out examples it never tuned on, and block the ship if it doesn't generalize. Sits on top of promptfoo/DSPy. C…

34
9

Clawdmarket

Autonomous agent-to-agent marketplace with live Karpathy loop self-improvement. Agents discover, hire, benchmark, and evolve programmatically. MPP/x402/MCP. No humans in the loop.

34
10

Humanloopbench

Human-in-the-loop LLM Failure Detection & Benchmark Toolkit — Auto-flag LLM failures (hallucination, refusal, factual error) with LLM-as-Judge explanations. Confirm them in a revi…

34
11

Ferryman MCP

A local-first MCP host with pluggable skills, multi-provider routing, and multi-channel I/O. Kotlin. Plus a Python eval harness with rule + LLM-judge scorers over a human-authored…

34
← Previous Page 1 of 1 Next →