Referee.Chat

Referee.Chat

Set a goal and a standard, a panel of AI models works it, and a referee rules on the evidence

Gallery

About Referee.Chat

Referee.Chat is a web app built on a simple frustration, which is that a single AI chat will confidently declare a problem solved whether or not it actually is. Here you set a goal, define the standard it has to meet, and then a panel of up to 6 AI models works through the problem together. A separate referee model, running on Claude Opus 4.5 or GPT-5, rules on each acceptance criterion individually, and it only accepts a criterion when there's verifiable evidence behind it. Opinions alone can't settle anything.

The evidence requirement has teeth. The system searches academic sources including arXiv, OpenAlex, Crossref, and Europe PMC, so claims get tied to sources that can be re-checked rather than to a model's recollection. It can also run Python code to verify computational claims directly, which turns arguments about whether a calculation holds into an executed check. When the panel gets stuck, the system hands the specific blocking question to the model best equipped for it instead of letting the whole panel thrash in circles.

Work is organized as matches, and a match behaves more like a project than a chat session. Matches pause and resume, and established findings get saved into a shared ledger the panel builds on rather than re-litigating. Costs are billed per token for the models you choose and scale with the size of that ledger, not with how many rounds the panel takes. You set a spending limit per match, and hitting it pauses the match rather than failing it, so you can review what's been established and raise the limit if the work justifies it. That's a meaningful safety property for long-running problems, because nothing you've paid for gets thrown away when the budget runs out, the work simply waits for you to come back and decide whether it's worth continuing.

How hard the referee pushes is up to you. Quality bars range from Careful Expert to Publication Standard, Formal Proof, and Prize Standard, so the same machinery can sanity-check an analysis or grind on something you'd want to defend in front of reviewers. The model roster spans Anthropic, OpenAI, Google, Mistral, DeepSeek, and others, so the panel isn't a single vendor debating itself, and you choose which models sit on the panel for a given match. There's also a documented API alongside the web interface, so the goal, panel, and referee process can be wired into your own workflow rather than lived in through the browser.

The division of labor inside a match is worth understanding, because it's the whole design. The panel models are the workers, arguing, drafting, and attacking the problem from different directions. The referee is deliberately separate and deliberately strong, running on Claude Opus 4.5 or GPT-5, and its job isn't to contribute ideas but to check them. A criterion doesn't clear because the panel agrees. It clears because there's a source that can be re-checked or a computation that actually ran. That's a different trust model from asking one model to both solve a problem and grade its own solution, which is what an ordinary chat quietly does.

The natural audience is researchers, mathematicians, and professionals whose decisions need to survive scrutiny. If you've used a chatbot for anything rigorous, you've likely met the failure mode this attacks, plausible-sounding claims resting on invented citations or unchecked arithmetic. For that audience the appeal isn't speed, it's that a ruling comes with the sources and executed checks behind it, so the result can be handed to a colleague, a reviewer, or your future self without asking anyone to take a chatbot's word for it.

Pricing is prepaid usage with no subscription. New accounts get a one-time $1.00 starting credit, enough to run a complete match on public settings. After that you top up from $10 in any amount, credit never expires, unused credit is refundable on request, and paid accounts unlock private matches and higher spending limits. There are no seats, no minimums, and no recurring fees. If you'd rather pay your model provider directly, you can bring your own inference API key, in which case Referee.Chat charges a 5% infrastructure fee that only kicks in after your first $5.00 of usage, with private matches included.

It's a young product with an unusual shape, closer to a structured review process than a chat tool, and the billing model matches that. You pay for the work done, the spend cap protects you mid-match, and nothing renews in the background. For questions where being wrong is expensive, that's a sensible way to buy AI effort.

Key Features

  • Panel of up to 6 AI models
  • Referee rulings on each criterion
  • Academic source search across arXiv and OpenAlex
  • Python execution to verify claims
  • Pausable matches with a shared ledger
  • Prepaid credit or bring your own key

Pros & Cons

What we like

  • Acceptance requires re-checkable evidence, not consensus
  • Spending caps pause a match instead of failing it
  • Credit never expires and is refundable on request
  • Multi-vendor model roster avoids one model grading itself

Room for improvement

  • No subscription tier, so heavy use means frequent top-ups
  • Free credit is a one-time $1 for public matches only
  • Rigorous matches on frontier models can get expensive
  • Younger product with a narrow, specialized focus

Frequently Asked Questions

What is Referee.Chat?
Referee.Chat is a web app where you set a goal and a standard, a panel of up to 6 AI models works the problem, and a referee model running on Claude Opus 4.5 or GPT-5 rules on each criterion using verifiable evidence from academic search and executed Python code.
Is Referee.Chat free?
New accounts get a one-time $1.00 starting credit, enough for a complete public match. After that it's prepaid credit topped up from $10, with no subscription, no seats, and no minimum. Credit never expires and unused credit is refundable on request.
How is Referee.Chat different from a normal AI chat?
A normal chat gives you one model's opinion. Here the work and the judging are separated. A multi-vendor panel does the work, and a referee only accepts criteria backed by re-checkable sources from arXiv, OpenAlex, Crossref, and Europe PMC, or by code that actually ran.
Can I use my own API key?
Yes. You can bring your own inference API key and pay your provider directly. Referee.Chat then charges a 5% infrastructure fee that starts only after your first $5.00 of usage, and private matches are included.

Best For

Stress-testing a research claim before publicationWorking a mathematical argument toward formal-proof standardVerifying computational results with executed codeComparing how far different frontier models get on one problem

Featured in

Alternatives to Referee.Chat

View all

Reviews (0)

No reviews yet

Be the first to share your experience with Referee.Chat

Sign in to write a review

Badge builder

Add Referee.Chat to your website

Choose a badge style and size, preview it here, then copy the generated HTML. Badge images are self-contained SVGs and do not require an external script.

Referee.Chat badge preview
<a href="https://toolindex.net/tools/referee-chat?ref=badge" target="_blank" rel="noopener">
  <img src="https://toolindex.net/badge/referee-chat/medium.svg" alt="Referee.Chat - Listed on Tool Index" width="180" height="50" />
</a>

How to use the badge

  1. 1. Pick the style, size, and theme that fit your layout.
  2. 2. Copy the generated HTML from the code block.
  3. 3. Paste it into your footer, homepage, or press page.

Standard badge available

The standard listing badge is available now. Score and circle badges are limited to tools currently ranked in the top 10 of a category.

Badge clicks return visitors to this profile with a referral tag so the source remains identifiable.