These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.
Coarena

Coarena is a free computer-use utility and live evaluation platform for frontier AI agents. A user enters one real browser or desktop task, and two leading agents attempt it in parallel with equivalent tools, sandboxes, budgets, and time limits. Users can inspect both trajectories, final answers, and downloadable artifacts, then choose A, B, Tie, or Both bad while model identities remain hidden. After the judgment, Coarena reveals the models and eligible blind preferences update a public leaderboard.
The platform supports continuously generated evaluation from real user work rather than relying only on a fixed benchmark. Its public benchmark defines 57 behavioral measures across task completion, speed, efficiency, recovery, cost, output quality, and human preference. Public governance documents the comparison protocol, eligibility rules, privacy treatment, anti-gaming safeguards, and rating methodology.
Coarena can help researchers, developers, policy teams, and AI users compare computer-use systems under consistent conditions, examine failures and recovery behavior, and understand how agent performance changes across realistic browsing, research, data, QA, and file-creation tasks.
About the tool
You can click on the links to see the associated tools
Tool type(s):
Objective(s):
Purpose(s):
Target sector(s):
Lifecycle stage(s):
Type of approach:
Usage rights:
Risk management stage(s):
Use Cases
Would you like to submit a use case for this tool?
If you have used this tool, we would love to know more about your experience.
Add use case




























