Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking
The startup aims to redefine how AI models are evaluated, targeting gaps in current systems. Founded in 2024, it seeks to establish a more accurate and reliable benchmarking framework.
Vals, a startup formed in 2024 and backed by Andreessen Horowitz, is positioning itself as a leader in AI benchmarking. The company is working to address shortcomings in existing systems, which have been criticized for being outdated and easily manipulated by AI firms.
Benchmarking has become a critical tool for AI companies to validate their models’ capabilities and differentiate themselves in a competitive market. However, many legacy systems have been shown to be vulnerable to manipulation, allowing companies to exaggerate their models’ performance.
In 2026, the industry is expected to face significant changes as new benchmarking standards emerge. Krishnan, a key figure at Vals, emphasized that the company’s approach focuses on evaluating the real-world impacts of AI models, rather than just theoretical performance metrics.
The shift toward more accurate benchmarking could increase costs for AI developers, as they may need to invest in more rigorous testing processes. It may also lead to greater vendor lock-in if companies rely heavily on specific benchmarking tools. Additionally, the change could influence governance practices and market reactions as stakeholders reassess AI capabilities.
Vals’ efforts are still in development, and the company acknowledges that its approach is not yet finalized. However, the startup’s mission to create a more reliable benchmarking system has drawn attention from industry observers and potential partners.