1. Conduid
  2. Marketplace
  3. #llm-evaluation
Tag

MCP servers tagged llm-evaluation

14 live MCP servers tagged "llm-evaluation", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with llm-evaluation.

14
live servers
48
avg trust
52
most stars
1

Opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

64
2

Trulens

Evaluation and Tracking for LLM Experiments and AI Agents

64
3

Agentic Security

Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪

64
4

Agenta

The open-source workspace for building and running AI agents. Build agents through chat, share them with your team, and run background agents on schedules or app events.

64
5

Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

59
6

Baba Is Eval

lennart-finke

Claude et al. play the brilliant puzzle title "Baba is You"

★ 52updated a year ago
56
7

Brain In The Fish

Score any document. Prove every claim.

44
8

Aicw AI Mentions

AICW AI Mentions (formerly AICW Rankings) is an open-source tool that helps online marketers track how often their products are mentioned in responses from popular AI chatbots.

39
9

Suede Creator Skills

25 open-source Agent Skills for Claude Code and Codex: multi-agent orchestration, A-F code review, AI evals, CI ship-gates, design, copy, SEO, iOS shipping, music rights, and cons…

39
10

Cap Evolve

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

39
11

Mcpscope

A local-first workbench for developing, inspecting, and benchmarking MCP servers against local (LM Studio, Ollama) or remote (OpenRouter) models — Web UI, CLI, and MCP interface o…

37
12

Rag Forge

Production-grade RAG pipelines with evaluation baked in

34
13

Ensemble

Multi-model consensus debate via the filesystem. LLMs propose, peer-review, rebut, vote and synthesize a group-confirmed answer. CLI + MCP.

34
14

Fde Guide

A field guide for FDEs and applied AI teams: value engineering and production architecture.

34
← Previous Page 1 of 1 Next →