Group of diverse individuals posing in a brick-walled room with plants, some seated and others standing.

Vals Bets That AI Benchmarking Will Become the New Trust Layer for Model Buyers

Vals is betting on AI benchmarking as the new trust layer, with a $40M round, rapid growth and private tests for real-world model performance.

In short

Vals, a San Francisco AI benchmarking startup backed by Andreessen Horowitz, says older tests are too easy to game and is building private, domain-specific evaluations for real-world model use. The company has raised $40 million, grown rapidly and is expanding into federal-agency work.

  • Vals raised a $40 million Series A led by Andreessen Horowitz.
  • The startup says public AI benchmarks are increasingly easy to game.
  • Vals focuses on private, domain-specific tests for real business tasks.
  • Revenue is up eightfold year over year, and the team has grown from 8 to 25.
  • The company has launched evaluation services for federal agencies.

Vals, the AI benchmarking startup backed by Andreessen Horowitz, is positioning itself as the industry’s next major standard-bearer for model evaluation after raising $40 million in a Series A and expanding rapidly through 2026. The San Francisco company argues that existing tests are too easy to game and too detached from how businesses actually use AI, making trustworthy benchmarking a critical piece of the market’s next phase.

Founded in 2024, Vals says it is building a more realistic way to test AI systems by focusing on domain-specific performance, hidden evaluation sets and the practical risks models could create in the real world. The company’s pitch matters because as AI buyers, investors and regulators demand clearer proof of capability, benchmarks are becoming more than a technical exercise: they are shaping sales, procurement and public trust.

At a time when companies routinely tout benchmark wins in marketing materials and investor decks, Vals is betting that the next wave of AI competition will hinge on who can measure models most credibly.

Why AI benchmarking has become such a contested market

AI benchmarking has turned into one of the most powerful tools in the industry because it offers a standardized way to compare models, validate claims and signal technical leadership. But it has also become a target for manipulation, as model developers increasingly optimize for the test rather than the real-world task.

That tension is now central to the market. Many of the older benchmark suites were built for a previous generation of machine learning systems and do not reflect how modern large language models behave across complex workflows. As models improve, they can overperform on public tests while still struggling in the actual jobs companies want them to do.

That gap is exactly what Vals says it is trying to close.

How benchmarks became marketing tools

Benchmarks once served primarily as technical yardsticks for researchers. Today, they are also a public relations instrument. When an AI lab posts a strong score, the result can become a headline, a social media victory lap and a selling point for enterprise customers.

The downside is that widely known tests can be reverse-engineered. If a benchmark is public, model developers may train against it directly, inflating scores without necessarily improving useful capability. That has made the credibility of evaluation a growing concern across the sector.

What Vals is building differently

Vals is building private evaluations designed to test what AI systems can actually do in professional settings, not simply what they know in the abstract. Instead of relying mainly on broad academic-style exams, the startup evaluates models on task execution in specific industries such as law, finance and software development.

The company’s approach centers on real-world outcomes: whether a model can produce work that matches human quality, where it fails, and what kinds of harm it might create if deployed at scale. That includes not only successful task completion but also negative behavior, misuse potential and broader safety implications.

In practice, that means Vals is trying to answer a question that has become increasingly important to enterprise buyers: can this model do the job, or can it only appear to do the job on a public leaderboard?

Rayan Krishnan, Vals’ co-founder, said the company was born from a belief that academic-style benchmarks had fallen behind the speed of frontier model development. He argued that as AI is woven into more parts of society, evaluations should verify whether products really deliver on the promises companies make about them.

Why secrecy matters to Vals

Unlike many benchmark providers that publish their test sets, Vals does not disclose its exact evaluation materials. The company believes that keeping tests hidden reduces the incentive for teams to train directly against them, which can distort results and weaken trust in the scores.

That secrecy is part of the company’s value proposition. In Vals’ view, a benchmark is only useful if it measures performance that has not already been optimized away by repeated exposure.

The startup also frames its work as more operational than academic. Rather than asking whether a model can answer trivia questions or pass generic exams, it asks whether the model can reliably produce usable output in the kinds of tasks businesses will pay to automate.

Who is behind Vals?

Vals was founded by Rayan Krishnan, a 25-year-old entrepreneur with experience at Palantir, Microsoft and Stanford’s well-known AI lab. His background reflects the blend of product, engineering and research perspectives that increasingly shape the AI startup ecosystem.

Krishnan has argued that benchmark design has not kept pace with model development. That concern helped inspire Vals’ founding in 2024, and the company has since moved quickly from an early-stage idea to a visible player in the evaluation market.

Last year, Vals secured a seed round led by 8VC and Bloomberg Beta. Then, after a period of fast growth, it raised a $40 million Series A led by Andreessen Horowitz. The funding round gives the startup more room to scale its testing infrastructure, expand its team and deepen its relationships with customers and government agencies.

How Vals is growing so quickly

Vals says revenue is now eight times higher than it was a year ago, a striking signal for a company that is still less than two years old. Its headcount has also climbed fast: the startup began 2026 with eight employees and has grown to 25, with plans to add another 10 to 15 people.

That growth suggests a rising willingness among AI companies to pay for independent evaluation services. It also reflects a broader market shift: as model spending increases, so does the cost of making bad deployment decisions. For enterprises, benchmarking is no longer just a research preference; it is becoming part of due diligence.

Vals is also preparing to move into a larger office as its staff expands. The company operates from a two-floor space on San Francisco’s Folsom Street, in a brick building that once housed a brewery. Today, the site is part of the city’s dense ecosystem of AI startups trying to turn technical expertise into durable businesses.

Milestone What happened Why it matters
2024 Vals was founded Startup begins with a focus on modern AI evaluation
Last year Seed round led by 8VC and Bloomberg Beta Early institutional backing validates the concept
Last month $40 million Series A led by Andreessen Horowitz Provides capital for expansion and product development
This year Team grows from 8 to 25 employees Signals rapid demand and operational scale-up
Recently Launch of federal-agency evaluation program Extends the company into public-sector AI assessment

What kinds of models does Vals test?

Vals evaluates models across both conventional business functions and more sensitive or unusual areas. Its portfolio includes assessments for law, finance and coding, but it also stretches into domains where consequences can be serious and errors costly.

The startup says it is working on evaluations related to recursive self-improvement, cybersecurity, biosecurity and mental health. It has also explored benchmark work tied to the law of armed conflict, including how AI systems might interpret or apply Geneva Convention-related principles.

That broader scope shows how the evaluation business is evolving. The industry is no longer just asking whether a model can write a paragraph, summarize a document or answer a test question. It is increasingly asking whether the model could operate safely in domains where mistakes can lead to financial losses, safety failures or legal exposure.

Why these domains matter

Specialized evaluations matter because the real value of AI is shifting from novelty to utility. A model that can ace a general benchmark may still be unreliable when asked to draft contracts, triage financial documents or support security workflows.

By focusing on high-stakes sectors, Vals is trying to become relevant not only to AI labs but also to buyers, compliance teams and public institutions that need a defensible basis for procurement decisions.

  • Law: Can the model handle legal reasoning and drafting accurately?
  • Finance: Can it process and interpret complex market or compliance tasks?
  • Coding: Can it complete practical software work at usable quality?
  • Security: Could it be misused in cyber or biosecurity contexts?

Why do companies pay to be tested?

Companies pay Vals because a negative result can still be valuable. If a model underperforms in a specific task, the vendor learns where to improve, and the buyer learns whether the system is ready for deployment.

Krishnan compares the model to standardized testing, suggesting that a company’s willingness to pay for evaluation is not unlike a student paying to take an exam. The test has value because it creates an external measure, even if the score is imperfect or disappointing.

That logic is increasingly relevant as enterprises sift through a crowded field of model providers. In many cases, the challenge is not finding an AI tool; it is knowing which one can be trusted in production.

Krishnan has said Vals is focused less on abstract measures of intelligence and more on the practical question of whether models can produce human-quality output in specific domains, while also revealing the risks they might pose if deployed broadly.

How federal agencies fit into the strategy

Vals recently launched a program focused on providing model evaluations to federal agencies, a move that could open an important new channel of demand. Government buyers tend to move more slowly than private companies, but they often need rigorous documentation, risk analysis and procurement-ready assessments.

That makes independent benchmark companies potentially attractive partners. Agencies evaluating AI use cases may need more than vendor claims; they may want third-party evidence that systems meet operational and safety requirements.

The public-sector work also gives Vals another way to position itself as a neutral arbiter rather than just another startup chasing enterprise contracts. In a field full of promotional claims, neutrality can become a product feature.

What Vals says about the future of AI trust

Vals believes benchmarking will become a central part of how AI companies raise money, win customers and describe their future plans. Krishnan has suggested that as AI businesses become more economically important and more publicly traded, the market will demand clearer evidence about model behavior and business risk.

That argument reflects a larger shift in the AI industry. In the early stages of the boom, capability demos were often enough to attract attention. Now, with AI embedded in products, workflows and infrastructure, investors and regulators are beginning to ask harder questions about reliability, safety and accountability.

In Krishnan’s view, future public filings, investor presentations and market disclosures may lean more heavily on benchmark and evaluation data than they do today. If that happens, companies like Vals could become the gatekeepers of credibility.

How this could reshape the market

If Vals succeeds, benchmarking could move from a back-end technical function into a core part of commercialization. Model providers would need to think not only about training and deployment, but also about how they will be measured by third parties and how those measurements will be read by customers and investors.

That would likely reward companies that build for durable performance rather than leaderboard optimization alone. It would also put pressure on benchmark designers to stay ahead of rapidly changing model behavior.

In other words, Vals is not just selling tests. It is trying to build the infrastructure for trust in an industry where trust is becoming expensive.

Timeline of Vals’ rise

The company’s ascent has been unusually fast for a startup in a niche infrastructure category. The following timeline captures the main steps in its growth so far.

  1. 2024: Vals is founded to modernize AI evaluation.
  2. 2025: The startup raises seed funding from 8VC and Bloomberg Beta.
  3. 2026, early: The team enters the year with eight employees.
  4. 2026, last month: Andreessen Horowitz leads a $40 million Series A.
  5. 2026, this month: The company says revenue is up eightfold year over year.
  6. 2026, recently: Vals launches a federal-agency evaluation program and begins planning further hiring.

What to watch next

Vals now faces the challenge that confronts many fast-growing infrastructure startups: proving that its approach becomes the standard others adopt rather than just one more option in a crowded market. To keep expanding, the company will need to show that its private benchmarks are both credible and commercially useful.

It will also need to manage the tension between secrecy and transparency. The more hidden the tests, the harder it may be for outsiders to validate them. The more public the process, the easier it becomes for models to game the system. Finding the right balance may determine whether Vals becomes a benchmark leader or simply another participant in an arms race.

For now, the company’s rapid fundraising, expanding staff and entry into government work indicate that investors and customers see real promise in the idea. In a sector obsessed with model performance, the company is betting that whoever controls the evaluation layer may eventually control much of the conversation around AI quality itself.

Frequently asked questions

What does Vals do in AI?

Vals builds AI benchmarking and evaluation tools that test whether models can perform real-world tasks in areas like law, finance and coding. The company says its goal is to measure practical usefulness, not just abstract intelligence or leaderboard performance.

Why are AI benchmarks a big deal now?

AI benchmarks are a big deal now because they influence buying decisions, investor perceptions and public trust. As models improve and public tests become easier to game, companies need more credible ways to prove that their systems work in production.

Who founded Vals?

Vals was founded by Rayan Krishnan, a 25-year-old entrepreneur with experience at Palantir, Microsoft and Stanford’s AI lab. He says the startup was created to address the gap between fast-moving frontier models and outdated academic evaluations.

How much funding has Vals raised?

Vals has raised a seed round and then a $40 million Series A led by Andreessen Horowitz. The new funding follows rapid revenue growth and gives the startup more room to hire, expand its office and deepen public-sector work.

Why does Vals keep its benchmarks private?

Vals keeps its benchmark materials private to reduce the chance that AI developers can train directly against the tests. The company believes hidden evaluations are more likely to reveal how models really perform instead of rewarding optimization for known questions.

Share this 🚀