Arena Raises $200M for AI Evaluation
Arena has raised $200M at a $3.1B valuation as AI evaluation becomes core enterprise infrastructure.

AI evaluation is becoming infrastructure rather than a side feature.
What happened
Arena raised a $200 million Series B at a $3.1 billion valuation.
The company builds systems for evaluating AI models and agents and is expanding into tools designed to measure whether agents act without authorisation, misrepresent task completion or behave outside intended constraints.
Arena says it has surpassed $100 million in annualised revenue.
Why it matters
As AI systems move from answering questions to taking actions, companies need ways to measure more than benchmark scores.
They need to know whether an agent follows instructions, behaves consistently and completes tasks safely in real workflows.
That creates a new enterprise-software category around independent evaluation, monitoring and alignment.
The bigger picture
AI deployment is creating an ecosystem around the models themselves.
Evaluation, observability, security and governance are becoming increasingly valuable because organisations cannot rely only on model providers to assess their own systems.
Arena's valuation suggests investors see that supporting layer as a large standalone market.
