Vitra

The $40M Bet on AI Evaluation: Why a16z Is Missing the Decentralized Trust Layer

Products | ProPanda |

The venture capital world has a peculiar habit of pouring money into infrastructure before the infrastructure is even needed. On Tuesday, a16z led a $40 million Series A round into Vals AI, a startup building evaluation tools for large language models. The pitch is straightforward: as enterprises rush to deploy AI, they need reliable, third-party tools to measure whether those models actually work—accuracy, safety, robustness. It’s a classic enterprise SaaS play, wrapped in the jargon of “AI governance” and “responsible deployment.”

But here’s the rub. The entire AI evaluation stack today is centralized. Vals AI’s tools run on their own servers, their own datasets, their own judgment algorithms. That means the same gatekeepers who decide what “good” looks like are also the ones profiting from the certification. Sound familiar? In blockchain, we call this a trust-minimization problem. And trust, as we learned in 2017, is the most expensive commodity in crypto.

Let me step back. I’ve spent the last six years in Web3, first as a protocol analyst during DeFi Summer, then as a community architect at Aave, and now building a DAO that bridges institutional capital with decentralized infrastructure. In that time, I’ve seen the same pattern repeat: a centralized solution gets funded, scales quickly, then hits a wall when trust breaks down. FTX is the obvious example. But the principle applies to any system where a single party controls the validation layer.

Vals AI claims to solve a real problem. According to their announcement, enterprises are struggling to move from AI proof-of-concept to production because they can’t trust model outputs. Without standardized evaluation, every deployment is a leap of faith. Vals AI’s platform promises to provide that evaluation—test suites, benchmarking, red-teaming tools—all in a tidy dashboard. The a16z investment signals that the market sees evaluation as the next “must-have” layer in the AI stack.

The $40M Bet on AI Evaluation: Why a16z Is Missing the Decentralized Trust Layer

But here’s where the blockchain lens changes everything. The evaluation of a model is itself a decision that can be gamed, biased, or corrupted. If Vals AI holds the keys to the evaluation dataset and the scoring logic, then any enterprise relying on their report is ultimately trusting Vals AI’s integrity. That’s a single point of failure. And in a world where AI models are becoming the backbone of financial systems, healthcare decisions, and legal reasoning, that level of trust is unsustainable.

This is not a theoretical critique. I’ve seen it happen in smart contract auditing: a centralized auditor gives a pass, a bug slips through, and millions get drained. The same will happen with AI evaluation if the process isn’t verifiable by a decentralized community. The industry needs something like a blockchain-based evaluation oracle—where the test data, the scoring algorithm, and the results are all on-chain, auditable, and subject to economic incentives that reward honest reporting.

The core insight is this: evaluation is a trust game, and the only way to minimize trust in a trust game is to decentralize the referee.

Let’s look at the technical landscape. Current AI evaluation tools like LangSmith, Galileo, and Arthur AI are all centralized. They decide what to test (e.g., hallucination rate, instruction following, bias) and how to score it (often using LLM-as-judge, which is another model that can be biased). The process is opaque—you can’t verify the evaluation protocol unless you trust the provider. Vals AI’s new product, which they claim is a “next-generation evaluation platform,” likely follows the same architecture.

But there are early signs of a decentralized alternative. Projects like Bittensor, Gensyn, and Ritual are building networks where validation is distributed across nodes, with cryptographic proofs of work. For example, Bittensor’s subnet for model evaluation rewards miners for producing accurate rankings, and the consensus mechanism ensures that no single actor can manipulate the results. This is closer to the ideal: a permissionless, transparent, and economically robust evaluation layer.

From my own experience working with institutional clients at Deutsche Bank’s digital assets desk, I’ve seen the demand for auditable AI. The bankers don’t just want a report—they want the ability to reproduce the evaluation themselves. They want to see the raw data, the model weights, the scoring logic. They want proof that the AI hasn’t been gamed. In traditional finance, that’s called regulatory compliance. In Web3, it’s called trustlessness.

Vals AI’s $40M round is a bet on centralized evaluation. But the real opportunity is in decentralized evaluation.

Consider the contrarian perspective: maybe centralized evaluation is good enough for now. After all, enterprises are still in the early stages of AI adoption. They need a simple, fast, and reliable tool to get started. And Vals AI, backed by a16z’s network, can rapidly onboard hundreds of clients. The quality of evaluation might be better than any decentralized alternative because a dedicated team can iterate faster and maintain higher data quality. The a16z thesis is that the market will reward speed and accuracy over decentralization.

That argument has merit. In my years building DeFi protocols, I’ve seen many idealistic projects fail because they prioritized decentralization over user experience. The market rewards what works, not what is philosophically pure. Vals AI could very well become the “Stripe for AI evaluation”—a centralized service that solves a real pain point and captures massive value. And a16z has a track record of backing such winners.

But here’s the blind spot: the very nature of AI evaluation is a coordination problem. The “correct” answer depends on context, values, and purpose. A centralized evaluator imposes a single set of values. That’s fine for low-stakes applications, but for high-stakes decisions—like credit scoring, medical diagnosis, or autonomous driving—the evaluation must be decentralized to avoid systematic bias. The EU AI Act already hints at this, requiring “human oversight” and “transparency” that point toward multi-stakeholder verification.

Moreover, the technology for decentralized evaluation is advancing rapidly. Zero-knowledge proofs can now verify that a model was evaluated on a specific dataset without revealing the data. Trusted execution environments can run evaluation within secure enclaves. On-chain reputation systems can reward honest evaluators and slash dishonest ones. The pieces are there; the missing link is a product that packages them into a seamless experience.

a16z missed the decentralized evaluation play because it is still a theoretical concept for most VCs. But the community is already building it.

I’ve been part of that community. In 2022, after the FTX collapse, I founded Resilience DAO to support displaced Web3 workers. We set up a mentorship network where senior developers helped juniors rebuild their careers. The lesson was clear: trust is not built by a single authority, but by a network of aligned individuals. The same principle applies to AI evaluation. The most resilient evaluation system will be one where no single actor can corrupt the results.

The takeaway is this: Vals AI’s funding is a validation of the evaluation market, but it is also a warning. If we let centralized players control the referee, we will repeat the same mistakes we made in DeFi—where oracles were centralized, and the entire system collapsed when they failed.

The next step is to build a decentralized evaluation layer that is composable with existing AI stacks. Imagine a protocol where you can submit your model, and a network of anonymous evaluators runs test suites, with results published on-chain. The cost might be higher today, but the trust premium is worth it. As the market matures, enterprises will demand auditability, not just speed. And when that happens, the decentralized evaluators will be ready.

Community is the only chain that cannot be broken.

In the end, the $40 million is a signal. It tells us that evaluation is the next frontier. But the frontier is not just about building better tools—it’s about building tools that anyone can verify. That’s the blockchain promise. And that’s the path I’ll be watching.

Market Prices

BTC Bitcoin
$77,781.1 +0.17%
ETH Ethereum
$2,404.79 -0.63%
SOL Solana
$100.89 +0.30%
BNB BNB Chain
$692.6 +0.58%
XRP XRP Ledger
$1.37 +0.86%
DOGE Dogecoin
$0.0830 +1.69%
ADA Cardano
$0.2051 +3.22%
AVAX Avalanche
$7.27 +0.55%
DOT Polkadot
$0.8753 -1.52%
LINK Chainlink
$11.19 -0.68%

Fear & Greed

65

Greed

Market Sentiment

Event Calendar

{{年份}}
15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

12
05
halving BCH Halving

Block reward halving event

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

28
03
unlock Arbitrum Token Unlock

92 million ARB released

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

18
03
unlock Sui Token Unlock

Team and early investor shares released

Altseason Index

41

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
# Coin Price
1
Bitcoin BTC
$77,781.1
1
Ethereum ETH
$2,404.79
1
Solana SOL
$100.89
1
BNB Chain BNB
$692.6
1
XRP Ledger XRP
$1.37
1
Dogecoin DOGE
$0.0830
1
Cardano ADA
$0.2051
1
Avalanche AVAX
$7.27
1
Polkadot DOT
$0.8753
1
Chainlink LINK
$11.19

🐋 Whale Tracker

🔴
0x396c...f067
6h ago
Out
4,296 ETH
🔵
0xe991...7235
3h ago
Stake
1,377,729 DOGE
🔵
0x410c...81d1
30m ago
Stake
2,386.30 BTC

💡 Smart Money

0x86f2...95f3
Experienced On-chain Trader
+$3.2M
78%
0x9f13...2180
Early Investor
+$4.8M
90%
0xaa80...93aa
Early Investor
+$0.7M
79%

Tools

All →