AI Benchmarking Startup Vals Wants to Become the Independent Scorekeeper for Models
Vals is building private, industry-specific evaluations for AI models as businesses look beyond public leaderboards to determine what systems can actually do in production.
AI model launches increasingly arrive with a wall of benchmark charts. The numbers can be useful, but they are also becoming harder to interpret as models improve and companies optimize around well-known tests.
TechCrunch reported on September 19 that Vals, backed by Andreessen Horowitz, is building an independent AI evaluation business designed to measure performance on practical tasks across industries such as law, finance and coding.
As AI becomes a core business dependency, independent evaluation is turning from a research exercise into infrastructure.
Public benchmarks have limits
A benchmark is most valuable when it measures a capability the model has not specifically trained to reproduce. Once a test is public and widely discussed, contamination and optimization become concerns.
A model can also score well on a general reasoning test while performing poorly on a company’s real workflow. A legal team, coding platform and customer-support operation care about very different failure modes.
Vals says it keeps specific test materials private and focuses on whether models can complete work that resembles real professional tasks.
Evaluation becomes procurement
For enterprises, model selection is increasingly a purchasing decision. Teams need evidence that one system is more reliable, cheaper or safer for a particular workload.
A serious evaluation may measure accuracy, latency, tool use, structured-output reliability, cost per successful task and the frequency of severe failures. Those results can influence which model receives production traffic.
The best model is becoming workload-specific rather than universal.
That is why internal evaluation suites are also becoming common. Developers can create a set of representative tasks from their own product, run every candidate model against the same tests and compare results before changing production traffic.
The trust business
Vals’ opportunity depends on credibility. An independent evaluator needs methodologies that customers understand and providers cannot easily game.
The company also has to evolve quickly. AI capabilities can change in months, while static academic benchmarks may take far longer to design and publish.
What comes next
Model evaluation is likely to become part of continuous deployment. Instead of evaluating an AI system once, companies may rerun tests every time a provider releases a new model or changes behavior.
That would make evaluation platforms similar to automated software testing: a permanent layer between development and production.
In a market overflowing with AI performance claims, trustworthy measurement can become a product of its own.