- Conduid
- Marketplace
- #regression-testing
MCP servers tagged regression-testing
19 live MCP servers tagged "regression-testing", ranked by trust score. Tags come from package metadata and repository topics, so this list covers servers across every category that work with regression-testing.
Proof
MCP server verification toolkit — wire-level protocol conformance, security hygiene, and record/replay contract regression, shipped as a client-ready delivery report plus a CI gat…
Coder Eval
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Contract
Contract testing for MCP servers: snapshot your tool surface, then fail CI when a change would break the agents that depend on it. Classifies breaking vs additive vs routing (desc…
redline
Zero-infrastructure QA agent — k6 performance + Playwright functional testing with red/green baselines and a deterministic, model-free deploy gate. Clone and run.
Vitest MCP Assert
Vitest integration for mcp-assert: run MCP server assertion YAML files as Vitest tests
Classifier Evals MCP Server
MCP server exposing classifier evaluation tools
Mcpward
Black-box security & contract testing for MCP servers — rug-pull, tool-poisoning, schema-drift & protocol checks for CI (JUnit + SARIF).
Livekit Agent Simulator
Black-box LiveKit voice agent testing: real WebRTC/SIP calls with an AI caller. Catch barge-in, noise, quiet-mic, and latency failures that text pytest misses. Forensic reports +…
Pinnedai
Permanent guardrails for AI-coded apps
Agent Prod
Production AI agent quality gate and risk control framework for LLMOps, agent evaluation, regression detection, gray release, audit, and observability
Behavior
Parity tests for MCP migrations, including tool results and observable side effects.
Claude Code Canary
Regression-test Claude Code releases, plugins, MCP servers and configs with deterministic scenarios, version bisecting and GitHub Actions.
RaceSimQA
QA regression testing tool for 'racing car' type of simulations. Analytics, Data visualisation, LLM chat and AI summary
Eval Harness For Solo Devs
Lightweight CLI + MCP server that lets solo devs run markdown‑based regression tests on any agent, catching silent breaks with instant diffs, cost & latency feedback.
TestTrout
🐟 The testing assistant for coding agents. Reads your repo, connects to your deployment, ranks what's untested, writes real tests, runs them — driven over MCP.
J Rig Skill Binary Eval
Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, tri…
Self Test
Automated regression testing for Flutter through direct callback invocation. Test your Flutter apps with AI agents (Claude) using 60+ Playwright-equivalent tools. Works on iOS, An…
Webtest Orch
Token-efficient e2e orchestration skill for Claude Code: explore once, replay deterministically. Playwright + axe-core + run-diff. Tests stay in your repo. MIT.
Agent Invariants
Deterministic behavior contracts for AI agents — preserve approvals, stop semantics, tool scope, recovery, and proof before done.