Ly Gravity

Vals AI’s $40M Raise: The Third-Party Evaluation Track Gets a Crypto-Style Reality Check

CryptoBear Weekly

Hook:

A 4,000-line GitHub PR. A hidden test suite. A machine grading a model’s output against real developer intent. That’s the pitch behind Vals AI’s $40 million Series A, led by a16z, at a $400 million valuation. The narrative is seductive: move beyond static benchmarks like GSM8K and HumanEval—both widely suspected of being contaminated by training data—and measure models on actual, messy engineering tasks. But speed readers beware: the same kinetic energy that makes this deal feel like a land grab also masks the structural risks that any forensic observer would flag. I don’t see a clear moat yet. I see a product that’s betting on opacity as a service.

Context:

Vals AI’s core innovation is not a new model architecture. It’s a evaluation infrastructure: it extracts pull requests from any GitHub repository, constructs hidden tests based on the original developer’s intent, and automatically judges whether the model’s code passes. This is essentially SWE-bench turned into a SaaS product. The company claims it covers finance, law, and medical domains, aiming to verify “production readiness” beyond code. The market timing is sharp: AI model vendors are desperate for third-party validation after public benchmarks lost credibility. But the article—sourced from a blockchain monitoring channel called “Dongcha Beating”—is a classic case of high signal, low verification. The funding event is real; the product claims are not.

Core:

Let’s cut through the hype with data. The $40 million Series A represents roughly 9–10% dilution at a $400 million valuation, which is standard for a Series A. But the valuation itself is a bet on a new category, not on current revenue. The company’s revenue claim— “this year’s revenue has reached 8 times the full-year 2025 expectation”—is a textbook example of ambiguous PR. The statement contains a timeline contradiction (2025 is not over yet), and the base number is undisclosed. Without ARR or customer count, that 8x figure is noise.

From my own experience auditing smart contract vulnerabilities during the DeFi summer, I’ve learned to distrust any evaluation system that doesn’t publish its test set. Vals uses historical PRs, which are publicly visible on GitHub. If those PRs were part of the training data for models like GPT-4 or Claude—which they almost certainly were, given the common crawl data used—then the evaluation is circular. The company claims to mitigate this by using private repositories for paying customers, but that’s a feature, not a guarantee. The question is: how does Vals prevent model vendors from reverse-engineering the hidden tests? The article doesn’t answer this. The technical risk is swept under the rug.

Furthermore, the claim that OpenAI, Anthropic, Google, Meta, and xAI cite Vals’ evaluations in their model cards is unverifiable. The analysis rates this as confidence C (low). Even if true, it’s a form of “paid endorsement” if those companies are also customers. The independence of a third-party evaluator that charges the very entities it evaluates is a conflict of interest that would get a traditional audit firm disqualified. In crypto, we call that a “token sale with no vesting.”

Contrarian:

The mainstream narrative is that Vals is the new standard for AI evaluation. The contrarian view: it’s a high-risk infrastructure play that faces three structural blind spots. First, the evaluation tasks are drawn from a narrow slice of software engineering—mostly open-source repos. Enterprise-specific workflows, especially in regulated industries like finance and healthcare, require custom annotation that Vals likely outsources. The article doesn’t disclose the human labor cost, which could be its biggest expense. Second, the “third-party” label is weakened by a16z’s involvement. a16z backs dozens of AI companies that could become Vals customers, turning the evaluator into a portfolio service. That’s not independent; it’s a network effect with a price tag. Third, the global impact is limited. Will Chinese AI labs or open-source models submit to a US VC-backed evaluation firm? Unlikely. This caps the market.

I see a parallel with the early days of smart contract auditing. Every DeFi protocol claimed to be “audited by CertiK,” but the audits were often shallow, and the auditors were paid by the clients. The market eventually demanded separate, conflict-free standards. Vals is entering that same phase, but with even less transparency. The company’s “hidden tests” are a black box. Without a public, verifiable audit trail, the evaluation results are just marketing.

Takeaway:

Vals AI’s $40 million raise is a signal that the AI industry is ready to commoditize evaluation. But the event is more about capital allocation than product maturity. The real test will come when customers demand to see the test set, or when a model passes Vals’ evaluation but fails in production. Then we’ll see if the infrastructure is a fortress or a facade. Watch for two things: the release of a public evaluation dataset from Vals, and whether any regulated financial institution signs a contract. Until then, this is a bet on a category, not a company.

I don’t think the valuation is justified by the disclosed metrics. I think the market is pricing narrative, not substance. I’ve seen this movie before—in 2017 ICOs, in 2020 DeFi farming, and now in AI evaluation. The music is playing, but the floor is still being built.

Market Prices

BTC Bitcoin
$76,638.8 -1.93%
ETH Ethereum
$2,379.53 -3.34%
SOL Solana
$97.95 -4.37%
BNB BNB Chain
$683.9 -0.55%
XRP XRP Ledger
$1.32 -4.58%
DOGE Dogecoin
$0.0810 -2.48%
ADA Cardano
$0.1942 -2.75%
AVAX Avalanche
$7.12 -2.25%
DOT Polkadot
$0.8444 -2.93%
LINK Chainlink
$11.02 -4.05%

Fear & Greed

63

Greed

Market Sentiment

Event Calendar

{{年份}}
12
05
halving BCH Halving

Block reward halving event

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

28
03
unlock Arbitrum Token Unlock

92 million ARB released

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

18
03
unlock Sui Token Unlock

Team and early investor shares released

Altseason Index

41

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
# Coin Price
1
Bitcoin BTC
$76,638.8
1
Ethereum ETH
$2,379.53
1
Solana SOL
$97.95
1
BNB Chain BNB
$683.9
1
XRP Ledger XRP
$1.32
1
Dogecoin DOGE
$0.0810
1
Cardano ADA
$0.1942
1
Avalanche AVAX
$7.12
1
Polkadot DOT
$0.8444
1
Chainlink LINK
$11.02

🐋 Whale Tracker

🔵
0x3e30...3cb2
1h ago
Stake
3,077,211 USDC
🔴
0xf3b8...b162
6h ago
Out
16,010 BNB
🔴
0xb59b...6c58
1h ago
Out
2,201,590 USDT

💡 Smart Money

0xe449...18f7
Early Investor
+$1.4M
67%
0x33d2...47d8
Top DeFi Miner
+$1.6M
95%
0xa10f...6471
Institutional Custody
+$4.6M
79%

Tools

All →