Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Free
Scanned where there is a package or repository to read, and every release diffed against the tool surface we already hold.
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.