The venture capital world has a peculiar habit of pouring money into infrastructure before the infrastructure is even needed. On Tuesday, a16z led a $40 million Series A round into Vals AI, a startup building evaluation tools for large language models. The pitch is straightforward: as enterprises rush to deploy AI, they need reliable, third-party tools to measure whether those models actually work—accuracy, safety, robustness. It’s a classic enterprise SaaS play, wrapped in the jargon of “AI governance” and “responsible deployment.”
But here’s the rub. The entire AI evaluation stack today is centralized. Vals AI’s tools run on their own servers, their own datasets, their own judgment algorithms. That means the same gatekeepers who decide what “good” looks like are also the ones profiting from the certification. Sound familiar? In blockchain, we call this a trust-minimization problem. And trust, as we learned in 2017, is the most expensive commodity in crypto.
Let me step back. I’ve spent the last six years in Web3, first as a protocol analyst during DeFi Summer, then as a community architect at Aave, and now building a DAO that bridges institutional capital with decentralized infrastructure. In that time, I’ve seen the same pattern repeat: a centralized solution gets funded, scales quickly, then hits a wall when trust breaks down. FTX is the obvious example. But the principle applies to any system where a single party controls the validation layer.
Vals AI claims to solve a real problem. According to their announcement, enterprises are struggling to move from AI proof-of-concept to production because they can’t trust model outputs. Without standardized evaluation, every deployment is a leap of faith. Vals AI’s platform promises to provide that evaluation—test suites, benchmarking, red-teaming tools—all in a tidy dashboard. The a16z investment signals that the market sees evaluation as the next “must-have” layer in the AI stack.

But here’s where the blockchain lens changes everything. The evaluation of a model is itself a decision that can be gamed, biased, or corrupted. If Vals AI holds the keys to the evaluation dataset and the scoring logic, then any enterprise relying on their report is ultimately trusting Vals AI’s integrity. That’s a single point of failure. And in a world where AI models are becoming the backbone of financial systems, healthcare decisions, and legal reasoning, that level of trust is unsustainable.
This is not a theoretical critique. I’ve seen it happen in smart contract auditing: a centralized auditor gives a pass, a bug slips through, and millions get drained. The same will happen with AI evaluation if the process isn’t verifiable by a decentralized community. The industry needs something like a blockchain-based evaluation oracle—where the test data, the scoring algorithm, and the results are all on-chain, auditable, and subject to economic incentives that reward honest reporting.
The core insight is this: evaluation is a trust game, and the only way to minimize trust in a trust game is to decentralize the referee.
Let’s look at the technical landscape. Current AI evaluation tools like LangSmith, Galileo, and Arthur AI are all centralized. They decide what to test (e.g., hallucination rate, instruction following, bias) and how to score it (often using LLM-as-judge, which is another model that can be biased). The process is opaque—you can’t verify the evaluation protocol unless you trust the provider. Vals AI’s new product, which they claim is a “next-generation evaluation platform,” likely follows the same architecture.
But there are early signs of a decentralized alternative. Projects like Bittensor, Gensyn, and Ritual are building networks where validation is distributed across nodes, with cryptographic proofs of work. For example, Bittensor’s subnet for model evaluation rewards miners for producing accurate rankings, and the consensus mechanism ensures that no single actor can manipulate the results. This is closer to the ideal: a permissionless, transparent, and economically robust evaluation layer.
From my own experience working with institutional clients at Deutsche Bank’s digital assets desk, I’ve seen the demand for auditable AI. The bankers don’t just want a report—they want the ability to reproduce the evaluation themselves. They want to see the raw data, the model weights, the scoring logic. They want proof that the AI hasn’t been gamed. In traditional finance, that’s called regulatory compliance. In Web3, it’s called trustlessness.
Vals AI’s $40M round is a bet on centralized evaluation. But the real opportunity is in decentralized evaluation.
Consider the contrarian perspective: maybe centralized evaluation is good enough for now. After all, enterprises are still in the early stages of AI adoption. They need a simple, fast, and reliable tool to get started. And Vals AI, backed by a16z’s network, can rapidly onboard hundreds of clients. The quality of evaluation might be better than any decentralized alternative because a dedicated team can iterate faster and maintain higher data quality. The a16z thesis is that the market will reward speed and accuracy over decentralization.
That argument has merit. In my years building DeFi protocols, I’ve seen many idealistic projects fail because they prioritized decentralization over user experience. The market rewards what works, not what is philosophically pure. Vals AI could very well become the “Stripe for AI evaluation”—a centralized service that solves a real pain point and captures massive value. And a16z has a track record of backing such winners.
But here’s the blind spot: the very nature of AI evaluation is a coordination problem. The “correct” answer depends on context, values, and purpose. A centralized evaluator imposes a single set of values. That’s fine for low-stakes applications, but for high-stakes decisions—like credit scoring, medical diagnosis, or autonomous driving—the evaluation must be decentralized to avoid systematic bias. The EU AI Act already hints at this, requiring “human oversight” and “transparency” that point toward multi-stakeholder verification.
Moreover, the technology for decentralized evaluation is advancing rapidly. Zero-knowledge proofs can now verify that a model was evaluated on a specific dataset without revealing the data. Trusted execution environments can run evaluation within secure enclaves. On-chain reputation systems can reward honest evaluators and slash dishonest ones. The pieces are there; the missing link is a product that packages them into a seamless experience.
a16z missed the decentralized evaluation play because it is still a theoretical concept for most VCs. But the community is already building it.
I’ve been part of that community. In 2022, after the FTX collapse, I founded Resilience DAO to support displaced Web3 workers. We set up a mentorship network where senior developers helped juniors rebuild their careers. The lesson was clear: trust is not built by a single authority, but by a network of aligned individuals. The same principle applies to AI evaluation. The most resilient evaluation system will be one where no single actor can corrupt the results.
The takeaway is this: Vals AI’s funding is a validation of the evaluation market, but it is also a warning. If we let centralized players control the referee, we will repeat the same mistakes we made in DeFi—where oracles were centralized, and the entire system collapsed when they failed.
The next step is to build a decentralized evaluation layer that is composable with existing AI stacks. Imagine a protocol where you can submit your model, and a network of anonymous evaluators runs test suites, with results published on-chain. The cost might be higher today, but the trust premium is worth it. As the market matures, enterprises will demand auditability, not just speed. And when that happens, the decentralized evaluators will be ready.
Community is the only chain that cannot be broken.
In the end, the $40 million is a signal. It tells us that evaluation is the next frontier. But the frontier is not just about building better tools—it’s about building tools that anyone can verify. That’s the blockchain promise. And that’s the path I’ll be watching.