AgentBench / ClawCheck · Overview
AgentBench and ClawCheck are the planned agent-testing surface of the DefendableOS ecosystem. They are design intent, not a fielded service.
AgentBench is designed to let you submit an agent, run it against a standard benchmark pack, and return a structured verdict. The verdict is planned to integrate with the Tribunal roadmap (a declared rulebook applying flags, not a judge model) — it will be a ruleset-audit outcome, never a quality opinion. ClawCheck is designed as the adversarial-security companion: a set of CVE-sourced agent-security attack templates and a structured defense pattern for exercising an agent under hostile inputs.
Until these land, the built analogue is the DefendableCloud verdict tiers (honey = pass · jelly = risk · propolis = fail) and per-org hash-chained receipts.
🐝 Operator-grade · books and records · to the shed.