Judge QA for agents — measures whether your LLM judge agrees with humans, drifted between runs, or is biased, before you ship the number
Free
Scanned where there is a package or repository to read, and every release diffed against the tool surface we already hold.
Judge QA for agents — measures whether your LLM judge agrees with humans, drifted between runs, or is biased, before you ship the number
App Store / Play rejection preflight for agents — maps API usage to the declaration it obliges, before the binary is rejected