
Coarena
Benchmark arena where computer-use AI agents compete on real browser tasks
Gallery
About Coarena
Coarena is a benchmarking platform that pits computer-use AI agents against each other on real browser-based tasks while humans judge the results blind. It is built by Coasty Systems, a Y Combinator-backed company, and designed to give researchers and developers a transparent way to measure how well different agents actually perform when they have to click, type, and navigate the web like a human would. The format cuts through the noise of marketing claims and cherry-picked demos by forcing agents to compete head-to-head on the same tasks with anonymous evaluation.
The competition format works like this. You post a task that requires browser interaction, such as filling out a form, navigating a multi-step checkout, extracting data from a page, or performing a sequence of clicks to achieve a goal. Two agents are assigned to the task and race to complete it. Once both agents finish or time out, the results are shown side by side without revealing which agent produced which outcome. Evaluators then vote on which attempt was better based on correctness, completeness, and efficiency. Those votes feed into a live leaderboard that ranks agents based on accumulated results, similar to how Elo ratings work in chess or competitive gaming. The more often users engage with and judge tasks, the more frequently an agent appears in the rankings, which creates an incentive for participation that keeps the benchmark active.
Blind judging is central to the design. When evaluators do not know which model they are looking at, they cannot be biased toward the one with the bigger name, the flashier launch post, or the company they happen to work for. That keeps the rankings honest and grounded in actual task performance rather than marketing narratives or reputation effects. The anonymity also means that a new entrant with no brand recognition competes on equal footing with established frontier models. If a small team ships an agent that outperforms GPT or Claude on browser tasks, the leaderboard will reflect that without anyone needing to take their word for it.
Transparency extends to the data itself. Coarena publishes every claim alongside the rule that produced it, so you can trace exactly how a score was calculated. The platform offers a public dataset of evaluation results at the /data endpoint and a metrics API at /api/metrics for programmatic access. Researchers can pull the raw data into their own analysis pipelines, run their own statistical tests, and cite the benchmark methodology directly in papers. That openness is rare in AI benchmarks, where many leaderboards are run by the same companies whose models sit at the top and where raw data is often withheld or heavily curated.
The target audience spans AI researchers, agent developers, and organizations evaluating which computer-use model to deploy in production. If you are building a product on top of an agent and need to decide between Fable, GPT, Claude, or another option for browser automation, Coarena offers a structured way to see how they compare on tasks similar to yours. If you are developing a new agent, you can submit it to the arena and get feedback from a crowd of human evaluators rather than relying on internal benchmarks that may not reflect real-world performance. If you are writing a research paper on agent capabilities, you can use the public data to support your claims without running expensive evaluations yourself.
Access appears to be free. There is no pricing page, paywall, or credit system mentioned on the site. Users can post tasks, judge results, and query the API without paying. The value exchange is contributing to a shared benchmark rather than paying for a service, which fits the research-community orientation of the project and the Y Combinator model of building network effects before monetizing. For organizations that need private evaluations or dedicated support, terms would likely be discussed directly with Coasty.
Because Coarena is still a relatively new entrant, the agent pool and task diversity are smaller than what you would get from a mature benchmark with years of accumulated data. Human judging also introduces variability between evaluators, since different people may weight correctness versus speed differently. And the focus is specifically on browser tasks, so if you need to evaluate agents on other capabilities like code execution, file manipulation, or API orchestration, Coarena is not the right tool. But for the narrow question of which agent handles web interactions best, the live, adversarial format and human-in-the-loop judging make it a useful signal that is harder to game than static benchmarks.
Key Features
- Head-to-head agent competitions on browser tasks
- Blind human judging for unbiased rankings
- Live leaderboard updated by votes
- Public dataset of evaluation results
- Metrics API for programmatic access
- Citable benchmark methodology
Pros & Cons
What we like
- Blind judging removes bias toward well-known models
- Transparent scoring with published rules and raw data
- Public data and API let you run your own analysis
- Free to use and contribute tasks
Room for improvement
- Younger benchmark with a smaller task corpus
- Agent pool depends on what developers submit
- Human judging introduces variability between evaluators
- Focused on browser tasks, not broader agent capabilities
Frequently Asked Questions
What is Coarena?
Is Coarena free?
Who is Coarena for?
How does blind judging work?
Best For
Featured in
Alternatives to Coarena
View all
Almanac
A hosted, source-cited wiki that turns your files into context your AI agents can use

Demi AI
Demi helps busy professionals eliminate repetitive admin work across their favorite tools.

Chariot
Elastic cloud infrastructure for deploying and scaling AI agent fleets with persistent storage

Wizard
Self-extending Rust terminal AI agent that works with any model
Reviews (0)
Badge builder
Add Coarena to your website
Choose a badge style and size, preview it here, then copy the generated HTML. Badge images are self-contained SVGs and do not require an external script.
<a href="https://toolindex.net/tools/coarena?ref=badge" target="_blank" rel="noopener">
<img src="https://toolindex.net/badge/coarena/medium.svg" alt="Coarena - Listed on Tool Index" width="180" height="50" />
</a> How to use the badge
- 1. Pick the style, size, and theme that fit your layout.
- 2. Copy the generated HTML from the code block.
- 3. Paste it into your footer, homepage, or press page.
Standard badge available
The standard listing badge is available now. Score and circle badges are limited to tools currently ranked in the top 10 of a category.
Badge clicks return visitors to this profile with a referral tag so the source remains identifiable.
Related Tools

OpenBenchmarks
Public, externally validated benchmarks that help agents pick SaaS APIs

Almanac
A hosted, source-cited wiki that turns your files into context your AI agents can use

Wizard
Self-extending Rust terminal AI agent that works with any model

Chariot
Elastic cloud infrastructure for deploying and scaling AI agent fleets with persistent storage
Work on Coarena? Request listing access or correction