About Humanloopbench
Humanloopbench is an MCP server in the AI category: human-in-the-loop LLM Failure Detection & Benchmark Toolkit — Auto-flag LLM failures (hallucination, refusal, factual error) with LLM-as-Judge explanations. Confirm them in a review UI or via MCP agent — then export a citable benchmark validated against expert labels. It has been installed 0 times through Conduid.
Install
git clone https://github.com/stevenincode/humanloopbenchThis server has no ConduID identity, so agent calls to it are not receipted. Pin the version you install and review the source before granting it credentials.
Ask AI
Ask AI about Humanloopbench
Powered by Claude · Grounded in docs
Security checks
- ·README presentNot checked yet.
- ·License declaredNot checked yet.
- ·Tests presentNot checked yet.
- ·Dependencies pinnedNot checked yet.
- ·No dynamic code executionNot checked yet.
- ·Scoped permissionsNot checked yet.