
Vals pushes to become AI's benchmark standard-bearer after $40M Series A
The AMW Read
Adds new baseline detail to yesterday's Series A story—federal-agency evals, private test methodology, safety-adjacent benchmarks—meaningful within data/eval infrastructure but not a structural cross-segment shift.
Vals pushes to become AI's benchmark standard-bearer after $40M Series A
Vals, an AI evaluation startup founded in 2024, is expanding after an a16z-led $40 million Series A last month, following an earlier seed round from 8VC and Bloomberg Beta. Co-founder Rayan Krishnan says academic benchmarks fell behind frontier model releases and became gameable once test materials leaked into training data. Vals keeps its tests private and grades models on completing real tasks in law, finance, and coding, plus newer benchmarks covering recursive self-improvement, mental health, cybersecurity, biosecurity, and Geneva Convention compliance. Revenue is now eight times last year's level, staff tripled from eight to 25 with 10-15 more hires planned, and the company just launched a program evaluating models for federal agencies.
The timing matters because benchmark credibility is shifting from a marketing signal to a governance one. Krishnan ties the push to Anthropic's expected listing this year and a possible OpenAI offering, arguing evaluation data could feed into public filings rather than just leaderboards. Per the AI Market Watch index, pipeline coverage of Vals rose from 76 to 111 tracked items over the trailing 90 days (name-matched, pipeline-ingested sources only), consistent with a company moving from a niche vendor toward broader attention.
A private, undisclosed test suite blocks train-on-the-exam gaming, but it also asks enterprise and government buyers to trust one vendor's methodology without a public audit trail — worth watching as Vals extends into biosecurity and armed-conflict-compliance testing. The more telling signal for investors is the federal contract and safety-adjacent benchmark expansion, not the round size: they position third-party evaluation as a recurring, regulatory-adjacent revenue line rather than a one-time diligence check.
#AIBenchmarking #ModelEvaluation #AndreessenHorowitz #AISafety #EnterpriseAI #Vals