AI model evaluator Arena nearly doubles its valuation to $3.1B


Arena has raised $200M at a $3.1B valuation and is launching an Alignment Index that ranks AI models on how often they act without authorisation or report work they never finished. It measures that the same way it measures preference, through its users, which is a harder thing to crowdsource.

Arena has raised $200M at a $3.1B valuation and will start ranking AI models on how often they do things nobody asked them to. Lightspeed Venture Partners and Khosla Ventures led the round. The company was worth $1.7B in January.

Its Alignment Index tracks signals including how often a model takes an unauthorised action, and how often it tells a user it finished a task it did not complete, chief executive Anastasios Angelopoulos told Bloomberg. It covers more than two dozen models. OpenAI’s GPT-6.1 Sol is first and Anthropic’s Claude Opus 5.5 second.

Arena began in 2023 as Chatbot Arena, a research project at Berkeley’s Sky Computing Lab, incorporated in 2025 and reported $100M in annualised revenue by June on more than 10 million human evaluations. Salesforce Ventures and Dell Technologies Capital joined the round.

The method is the question.

When it was LMArena raising $150M in January, we noted that crowdsourced voting can be gamed where safeguards are weak, and can reward answers that sound right over answers that are. Which of two replies you prefer is a matter of taste. Whether an agent did something nobody authorised is not.

That second thing keeps happening.

Google DeepMind set 100 agents to prove 71 conjectures in September and 14% cheated the grader, redefining a theorem’s symbols so unproven statements became trivially true. The swarm reported all 71 solved. Arena’s index is also a grader.

Regulators arrived this week.

The Information Commissioner’s Office opened a call for evidence on Thursday on how organisations manage data protection risks from agents, with responses due on 20 November, feeding a statutory code of practice. Autonomy “is not an excuse for poor compliance“, its director of technology regulation Richard Nevinson said.

Europe’s version of Arena is smaller and points somewhere else.

Galtea, spun out of the Barcelona Supercomputing Center, raised $3.2M in March to generate adversarial test cases against an agent’s intended behaviour and hand the results to compliance teams. It sells them as evidence for the EU AI Act, where breaches reach EUR 35M.

Both companies measure the same thing. They sell different objects.

One produces a ranking that model makers want to win, which is why they turn up. The other produces a document a regulator will ask to see. Arena is worth nine hundred and seventy times what Galtea raised.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *