Hook:
A 4,000-line GitHub PR. A hidden test suite. A machine grading a model’s output against real developer intent. That’s the pitch behind Vals AI’s $40 million Series A, led by a16z, at a $400 million valuation. The narrative is seductive: move beyond static benchmarks like GSM8K and HumanEval—both widely suspected of being contaminated by training data—and measure models on actual, messy engineering tasks. But speed readers beware: the same kinetic energy that makes this deal feel like a land grab also masks the structural risks that any forensic observer would flag. I don’t see a clear moat yet. I see a product that’s betting on opacity as a service.
Context:
Vals AI’s core innovation is not a new model architecture. It’s a evaluation infrastructure: it extracts pull requests from any GitHub repository, constructs hidden tests based on the original developer’s intent, and automatically judges whether the model’s code passes. This is essentially SWE-bench turned into a SaaS product. The company claims it covers finance, law, and medical domains, aiming to verify “production readiness” beyond code. The market timing is sharp: AI model vendors are desperate for third-party validation after public benchmarks lost credibility. But the article—sourced from a blockchain monitoring channel called “Dongcha Beating”—is a classic case of high signal, low verification. The funding event is real; the product claims are not.
Core:
Let’s cut through the hype with data. The $40 million Series A represents roughly 9–10% dilution at a $400 million valuation, which is standard for a Series A. But the valuation itself is a bet on a new category, not on current revenue. The company’s revenue claim— “this year’s revenue has reached 8 times the full-year 2025 expectation”—is a textbook example of ambiguous PR. The statement contains a timeline contradiction (2025 is not over yet), and the base number is undisclosed. Without ARR or customer count, that 8x figure is noise.
From my own experience auditing smart contract vulnerabilities during the DeFi summer, I’ve learned to distrust any evaluation system that doesn’t publish its test set. Vals uses historical PRs, which are publicly visible on GitHub. If those PRs were part of the training data for models like GPT-4 or Claude—which they almost certainly were, given the common crawl data used—then the evaluation is circular. The company claims to mitigate this by using private repositories for paying customers, but that’s a feature, not a guarantee. The question is: how does Vals prevent model vendors from reverse-engineering the hidden tests? The article doesn’t answer this. The technical risk is swept under the rug.
Furthermore, the claim that OpenAI, Anthropic, Google, Meta, and xAI cite Vals’ evaluations in their model cards is unverifiable. The analysis rates this as confidence C (low). Even if true, it’s a form of “paid endorsement” if those companies are also customers. The independence of a third-party evaluator that charges the very entities it evaluates is a conflict of interest that would get a traditional audit firm disqualified. In crypto, we call that a “token sale with no vesting.”
Contrarian:
The mainstream narrative is that Vals is the new standard for AI evaluation. The contrarian view: it’s a high-risk infrastructure play that faces three structural blind spots. First, the evaluation tasks are drawn from a narrow slice of software engineering—mostly open-source repos. Enterprise-specific workflows, especially in regulated industries like finance and healthcare, require custom annotation that Vals likely outsources. The article doesn’t disclose the human labor cost, which could be its biggest expense. Second, the “third-party” label is weakened by a16z’s involvement. a16z backs dozens of AI companies that could become Vals customers, turning the evaluator into a portfolio service. That’s not independent; it’s a network effect with a price tag. Third, the global impact is limited. Will Chinese AI labs or open-source models submit to a US VC-backed evaluation firm? Unlikely. This caps the market.
I see a parallel with the early days of smart contract auditing. Every DeFi protocol claimed to be “audited by CertiK,” but the audits were often shallow, and the auditors were paid by the clients. The market eventually demanded separate, conflict-free standards. Vals is entering that same phase, but with even less transparency. The company’s “hidden tests” are a black box. Without a public, verifiable audit trail, the evaluation results are just marketing.
Takeaway:
Vals AI’s $40 million raise is a signal that the AI industry is ready to commoditize evaluation. But the event is more about capital allocation than product maturity. The real test will come when customers demand to see the test set, or when a model passes Vals’ evaluation but fails in production. Then we’ll see if the infrastructure is a fortress or a facade. Watch for two things: the release of a public evaluation dataset from Vals, and whether any regulated financial institution signs a contract. Until then, this is a bet on a category, not a company.
— I don’t think the valuation is justified by the disclosed metrics. I think the market is pricing narrative, not substance. I’ve seen this movie before—in 2017 ICOs, in 2020 DeFi farming, and now in AI evaluation. The music is playing, but the floor is still being built.