Vals raises $40 million to expand confidential, real-task AI evaluations
Founded in 2024, Vals raised a $40 million Series A led by Andreessen Horowitz and says it will expand confidential evaluations built around real work in fields including law, finance, programming, cybersecurity, mental health and biosafety.

Vals Secures a $40 Million Series A to Redefine AI Benchmarking
In early 2024, Vals announced the closing of a Series A financing round that raised $40 million. The round was led by Andreessen Horowitz, following an earlier seed round that had been organized by 8VC and Bloomberg Beta. This infusion of capital marks a significant milestone for a company that was founded only months earlier, and it provides the financial backbone for the firm’s ambition to overhaul the way artificial‑intelligence systems are evaluated.
The core motivation behind Vals’ fundraising effort is to replace what the company describes as "old public benchmarks" with a new class of assessments. According to Vals, many existing benchmarks suffer from data contamination, meaning that the datasets used for evaluation may have already been incorporated into the training pipelines of leading models. By keeping the evaluation materials confidential, Vals aims to create a testing environment that is insulated from such leakage, thereby delivering a clearer picture of a model’s true capabilities.
From Abstract Exams to Real‑World Tasks
Rayan Krishnan, the founder of Vals, has articulated a specific philosophy regarding what constitutes a meaningful benchmark. He argues that a benchmark should verify whether a model can produce work of human‑level quality within a concrete domain, rather than merely succeeding at an abstract, multiple‑choice style exam. This perspective informs the selection of task categories that Vals currently measures, which include law, finance, programming, cybersecurity, mental health, biosafety, and the law of armed conflict.
Each of these domains represents a distinct set of professional standards and regulatory frameworks. By focusing on tasks such as drafting legal arguments, performing financial risk analysis, writing production‑grade code, identifying cyber threats, delivering mental‑health counseling scripts, assessing biosafety protocols, or interpreting the law of armed conflict, Vals seeks to capture performance dimensions that are directly relevant to end‑users and stakeholders.
Business Model and Market Adoption
Vals operates on a business‑to‑business model in which AI model developers pay the company to have their systems evaluated against its proprietary tasks. The resulting scores are then made available to prospective buyers, who can use them to compare competing systems on criteria that matter in real‑world deployments. This two‑sided market structure creates an incentive for model providers to demonstrate competence in Vals’ tasks while giving buyers a more nuanced decision‑making tool.
The company reports that its revenue has grown eightfold over the past year, a claim that aligns with the rapid scaling of its client base. Concurrently, Vals expanded its staff from eight to twenty‑five employees and announced plans to recruit an additional ten to fifteen people in the near future. These figures suggest a correlation between the influx of capital, the adoption of its evaluation services, and organizational growth.
Among the early adopters of Vals’ assessments are several United States federal agencies. The firm has launched a dedicated evaluation program for these agencies, indicating that its methodology is being considered for official procurement and compliance contexts. This development could signal a broader governmental interest in standardized, task‑based AI performance metrics.
Transparency, Independence, and Reproducibility
Vals publishes both the results of its evaluations and the underlying methodologies on its public website. This approach is intended to provide visibility into how scores are derived and to allow external parties to scrutinize the process. However, the company’s independence must still be examined, particularly because the entities that fund the evaluations are also the subjects of those evaluations.
The potential conflict of interest arises from the fact that model developers pay Vals for testing, which could influence the perceived impartiality of the outcomes. While Vals asserts that its methods are robust, the reproducibility of its results by independent researchers remains an open question, given that the evaluation materials themselves are kept confidential.
- Confidential test materials reduce risk of training‑data leakage
- Revenue model ties evaluation fees to model developers
- Results are shared with buyers for comparative analysis
- Public disclosure of methods aims to enhance transparency
The confidentiality of the test sets is a double‑edged sword. On one hand, it protects the integrity of the evaluation by preventing models from overfitting to known data. On the other hand, it limits the ability of third‑party auditors to replicate the tests, which could affect confidence in the reported scores.
Another area of uncertainty concerns the scalability of Vals’ task suite. While the current set of domains covers a broad spectrum of professional activities, extending the framework to additional sectors—such as education, transportation, or environmental monitoring—may require new task designs, subject‑matter expertise, and possibly distinct confidentiality protocols.
The company’s rapid staff expansion also raises questions about the sustainability of its operational model. Hiring ten to fifteen additional employees suggests a need for more evaluators, data engineers, and domain experts, but the long‑term demand for such specialized assessments will depend on how widely the industry adopts task‑based benchmarking.
Looking ahead, Vals’ approach could influence broader industry standards. If its confidential, task‑oriented assessments gain acceptance among both private firms and public agencies, they may become a reference point for future regulatory frameworks that seek to ensure AI systems meet concrete performance thresholds.
Nevertheless, the ultimate impact of Vals’ methodology will hinge on several unresolved factors: the ability of independent researchers to verify results without access to the test content, the willingness of model developers to continue paying for proprietary evaluations, and the degree to which buyers rely on these scores in procurement decisions. Each of these variables could affect whether Vals’ model becomes a dominant paradigm or remains a niche service for specialized applications.
Sources
- Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarkingTechCrunch · September 19, 2026
- Independent Evaluation, Unbiased BenchmarksVals AI · September 21, 2026



