Benchmarking has become the industry standard for how AI companies validate model functionality and stand out from competitors and tout their advantages when metrics tilt in their favor. In other words, good benchmarking almost always means good PR.
Unfortunately, companies have also figured out how to outsmart traditional benchmarking systems. Many of them are old and not built to measure the capabilities of modern models.
Vals, a startup founded in 2024, says its mission is to fix this very imperfect system. In less than two years, the company has established itself as a force to be reckoned with in the technology industry, and last year successfully secured a seed round led by 8VC and Bloomberg Beta. After rapid growth, the company raised $40 million in Series A led by Andreessen Horowitz last month.
The company’s co-founder, Rayanne Krishnan, 25, previously interned at Palantir and worked as an undergraduate at Stanford University at Microsoft and the school’s lauded artificial intelligence lab. Krishnan said Vals was born out of his own observations about how benchmarks were lagging behind the progress of the industries they were designed to measure.
“We’ve seen a number of very capable new models come to market rapidly, but academic benchmarks haven’t kept up with advances on that frontier,” says Krishnan. As AI is integrated into every part of society, Krishnan said there should actually be benchmarks to verify that models can do what companies advertise they can do.
Last week, the young founder gave me a tour of his company’s two-story office on Folsom Street in San Francisco. This old brick building was the site of a large brewery a century ago. This historic building is now no longer an industrial beer production site, but home to a variety of start-up companies looking to build the future of the technology industry.
“I think historically, assessments have been done to assess intelligence in a very abstract way,” Krishnan told me. “For example, does the model know enough information to take a bar exam type test?”
This is where Vals sets itself apart. While many benchmarking systems offer publicly available tests (which allows companies to train models against those tests and possibly cheat on exams), Vals does not publish its specific test materials. Instead of measuring the general knowledge of AI models, Vals also evaluates models on their ability to complete complex tasks related to specific industries such as law, finance, and coding.
“What we’re doing is actually looking at the real-world impact of the model,” Krishnan said. “Can we work to produce products of the same quality as humans in all areas?”
The idea is to check for negative results as well as positive ones, he says. What is expected is an analysis of “what kind of negative impact would there be if these models were to run out of control around the world?”
The capabilities Vals is measuring are growing. In addition to more traditional industries, the startup continues to expand into more unique areas. “We have benchmarks for reflexive self-improvement. We are working on models for understanding how to apply the Geneva Conventions in mental health, cybersecurity, biosecurity, and even the law of armed conflict,” says Krishnan.
Companies pay Vals to test their models, which can be a strange concept. Why would companies pay to find out that their models aren’t working? But having effective measurements can help companies troubleshoot and improve over time. Krishnan compares their revenue model to how students pay college boards to take the SAT.
Moreover, these evaluations have become important decision-making factors for companies looking to acquire new AI models.
The startup recently revealed that its revenue is now eight times higher than last year. The number of staff is also increasing. The Vals’ team, which started this year with just eight players, has already tripled to 25 players. As the startup grows, Krishnan said he plans to move into a larger office and hire 10 to 15 more people. The company also recently launched a program focused on providing model evaluations to federal agencies.
Krishnan believes his benchmarking system is the future of how AI companies think about growing their businesses and building public trust.
“AI companies are starting to go public. SpaceX went public, Anthropic is scheduled for later this year, and I think OpenAI will go public soon as well. As AI models become core to the economy and become more widespread, the types of benchmarks and evaluations we do will drive their use and become core to how these companies file public offerings and talk about future investments in AI.”
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
