Chasing the alpha through the digital fog – I’ve been watching the AI evaluation space since 2023, when I first noticed that the same protocols used to audit smart contracts were being retrofitted for large language models. Last week, a16z dropped a $40M Series A into Vals AI, a company that builds tools to measure how reliable AI models actually are. The news hit Crypto Briefing with the usual paucity of technical detail – no revenue numbers, no customer list, no architecture diagrams. But as someone who spent the 2017 ICO cycle reading Solidity code when everyone else was reading whitepapers, I know that a funding round like this is never just about the money. It’s a narrative signal. And in a sideways market, narrative is the only asset class that’s still appreciating.
Context: The invisible infrastructure of the AI-crypto convergence
Vals AI operates in what I call the “trust layer” of AI – the middleware that sits between raw model outputs and production deployment. The company’s pitch is simple: as enterprises rush to integrate LLMs into their workflows, they need a way to verify that the model isn’t hallucinating, leaking data, or making biased decisions. This is the same problem that blockchain solved for financial transactions, but now applied to software behavior. The $40M raise, led by a16z’s crypto and AI infrastructure teams, signals that the two worlds are converging faster than most analysts expect.

Based on my experience building the “DeFi Narrative Architect” series in 2020, I’ve learned that infrastructure waves follow a predictable pattern: first come the protocols, then the monitoring tools, then the governance frameworks. Vals AI is a second-wave play – the monitoring tool for AI. But the market is already crowded with players like LangSmith, Galileo, and Arthur AI. Why bet on Vals? The answer lies in the funding structure: a16z’s involvement suggests they see Vals as a platform, not just a point solution. The same firm that backed Coinbase and Uniswap is now backing an AI evaluation tool. That’s a narrative shift worth analyzing.

Core: The technical mechanics of evaluation as a service
Let me be clear: I haven’t audited Vals AI’s codebase. But I have audited enough AI evaluation pipelines to understand the landscape. The standard approach today is “LLM-as-Judge” – using one powerful model (like GPT-4o) to evaluate the outputs of another model. This is elegant but fragile. The judge model can be biased, inconsistent, or even gamed. Vals AI’s claimed innovation is a proprietary evaluation methodology that combines multiple judges, scenario-based testing, and automated red-teaming. The key question is whether they’ve solved the “who evaluates the evaluator” paradox.
Mapping the invisible architecture of value – what I find most interesting is the potential for blockchain-based verification of evaluation results. Imagine a future where an AI model’s performance is attested by a set of independent evaluators, and those attestations are stored on-chain via zero-knowledge proofs. This is exactly the kind of infrastructure that would make AI models auditable for regulatory compliance (like MiCA’s requirements for automated decision-making). Vals AI’s tools could become the backbone of that system – but only if they embrace cryptographic transparency.
Contrarian: The audit theater trap
Here’s the uncomfortable truth that the article glosses over: evaluation tools can create a false sense of security. During the NFT boom, I embedded myself in the Bored Ape Yacht Club Discord and interviewed 200 holders. I learned that status signaling often trumps substance. The same dynamic applies to AI evaluation. Companies will use a tool like Vals AI to generate a “safety score” and then present it to regulators without understanding the limitations of the methodology. This is “audit theater” – the same phenomenon that plagued DeFi audits in 2021, where a smart contract could pass a review but still get exploited.
Anthropology of the tokenized soul – the real value of Vals AI’s product may not be technical accuracy but narrative legitimacy. a16z is betting that enterprises will pay for the appearance of rigorous evaluation, even if the underlying methodology is imperfect. This is a cynical view, but it’s consistent with how markets work. The contrarian angle is that the evaluation tool market itself is a bubble, and the $40M raise is a peak signal. But I’ve seen this film before: in 2017, when Tezos raised $232M, everyone said the same thing. The code had flaws, but the narrative of self-amending ledger was powerful enough to sustain a multi-year ecosystem.
Takeaway: The next narrative is the fusion of evaluation and on-chain verification
The narrative is the new liquidity – Vals AI’s funding is not just a bet on AI infrastructure; it’s a bet on the convergence of two trust systems. Blockchain gave us tamper-proof transaction records. AI evaluation tools give us tamper-proof model performance records. The next step is to combine them. I see a future where every AI model deployed in a regulated industry is accompanied by an on-chain attestation of its evaluation results, verified by a DAO of independent evaluators. That’s the alpha hiding in the digital fog. Vals AI may be the first domino, but the game is just beginning.