The numbers landed in my terminal like a forensic anomaly. Vercel's CEO, Guillermo Rauch, published a dataset that should have shattered every narrative about the AI model market. Open-source models now account for 62% of all tokens processed on the platform. Yet they represent only 8.6% of the spending. The gap is not a rounding error. It is a structural revelation.
For anyone who has spent the last two years tracking on-chain data and protocol flows, this pattern is familiar. The market is pricing two different realities. The token share tells you where the volume is. The expenditure tells you where the value is. They are not the same thing. And the divergence is not a temporary blip. It is the new architecture of the AI economy.
I have spent the last decade building quantitative models for risk assessment. I have audited ICO whitepapers in 2017, stress-tested Uniswap V2 pools during DeFi Summer, and reverse-engineered the Terra collapse in 2022. The one lesson that persists across every market cycle is this: volume without value is a trap. The Vercel data is a textbook case.
Context: The Platform as a Neutral Observer
Vercel is not an AI company. It is a deployment platform for web developers. Its AI Gateway routes API calls to multiple model providers, making it a rare neutral observer in a market dominated by vendor marketing. When Rauch shared the data on August 22, 2024, he provided a cross-section of real developer behavior. This is not a survey. It is not a benchmark. It is a ledger of actual usage.
The dataset covers a period of rapid change. In early 2024, closed-source models dominated the token flow. OpenAI and Anthropic were the default choices for serious work. Google's Gemini was a distant third. Open-source models were a cost-saving alternative for low-stakes tasks. The split was roughly 72% closed, 28% open. By August, the ratio had inverted. Open-source models now process 62% of all tokens. The shift took less than six months.
This is not a marginal change. It is a migration. Developers have moved their workloads en masse. The question is why, and what it means for the industry's future.
Core: The Forensic Breakdown of the Data
Let me walk through the numbers with the precision they deserve. The dataset provides three key variables: token share, expenditure share, and provider ranking. Each tells a different story.
Token Share: The Volume Migration
Open-source models now handle 62% of all tokens on Vercel. This is a doubling from the 28.4% recorded earlier in the year. The growth is not linear. It is exponential. The inflection point appears to have occurred around the release of DeepSeek-V2 and the subsequent V3 series.
DeepSeek's architecture is the key variable. The mixture-of-experts (MoE) design combined with multi-head latent attention (MLA) delivers a significant reduction in inference cost. The token price is an order of magnitude lower than GPT-4o or Claude 3.5. This is not a subsidy. It is an engineering advantage. The cost structure is fundamentally different.
Developers are rational actors. They will not sacrifice quality for price if the output is unacceptable. The fact that they have migrated in such numbers indicates that open-source models have crossed a usability threshold for a significant portion of tasks. Code completion, simple refactoring, documentation generation, and test case writing are now handled effectively by open-source models. The quality gap has narrowed to the point where it is no longer a deciding factor for these workloads.
Expenditure Share: The Value Concentration
Here is the paradox. Open-source models process 62% of the tokens but account for only 8.6% of the spending. Closed-source models process 38% of the tokens but account for 91.4% of the expenditure. The unit economics are stark. The average price per token for open-source models is approximately 1/14th that of closed-source models.
This is not a reflection of cost. It is a reflection of pricing strategy. Open-source providers are using penetration pricing. They are trading margin for market share. The goal is to establish an ecosystem position before the market matures. This is a classic playbook, but the scale is unprecedented.
Anthropic is the most interesting data point. It processes 30% of the tokens but captures 65.1% of the expenditure. Claude 3.5 Sonnet is priced at $3 per million input tokens and $15 per million output tokens. This is higher than GPT-4o's $2.5/$10. Yet developers are paying the premium. The reason is capability. For complex code generation, long-document analysis, and agentic workflows, Claude is perceived as superior. The premium is a quality signal.
Provider Ranking: The DeepSeek Disruption
DeepSeek has surpassed Google to become the second-largest model provider on Vercel. This is a milestone. Google's Gemini series has been positioned as a top-tier competitor. The fact that it has been overtaken by a Chinese open-source model is a structural shift.
The implications are twofold. First, Google's developer ecosystem is weaker than its research capabilities suggest. The API pricing is not competitive. The iteration cycle is inconsistent. The developer tools are not sticky. Research leadership does not translate to product adoption. Second, DeepSeek's rise is not just about price. The model's performance on code and Chinese-language tasks has earned community trust. Developers are not just saving money. They are getting acceptable results.
The Long Tail Effect
The 62% token share likely includes a significant volume of low-value, high-frequency calls. Batch data processing, embedding generation, and simple classification tasks are ideal for open-source models. These are not the tasks that drive revenue. They are the tasks that drive volume. This explains the divergence between token share and expenditure share.
The high-value tasks—complex reasoning, multi-step tool use, and enterprise-grade code generation—remain the domain of closed-source models. This is where the money is. And this is where the competitive moat is deepest.
Contrarian: Correlation Is Not Causation
The obvious conclusion is that open-source models are winning. The token share suggests a mass migration. The expenditure share suggests a value gap. But this interpretation is incomplete. Correlation is not causation. The data is a snapshot, not a trend line.
First, the Vercel sample is biased. The platform serves web developers and front-end engineers. This is a specific segment of the AI market. It is not representative of enterprise deployments, financial services, or scientific computing. The token share for open-source models in these sectors is likely lower. The Vercel data may overstate the open-source penetration rate.
Second, the expenditure share does not capture the total cost of ownership. The 8.6% figure only reflects API call costs. It does not include the cost of self-hosting open-source models. GPU infrastructure, operational overhead, and engineering time are significant. For many organizations, the total cost of running an open-source model is higher than the API cost suggests. The apparent cost advantage may be an illusion.
Third, the token share is a lagging indicator. It reflects past decisions, not future intentions. The migration to open-source models may have been driven by a specific set of circumstances—the release of DeepSeek-V3, the pricing adjustments, the quality improvements. These factors may not persist. The market is dynamic. The next model release from OpenAI or Anthropic could shift the balance again.
Fourth, the value concentration in closed-source models is not a permanent state. It is a function of the current capability gap. If open-source models close the gap on complex reasoning and tool use, the expenditure share will shift. The 91.4% figure is not a law of nature. It is a snapshot of a competitive equilibrium that is under pressure.
The Structural Implications
The Vercel data reveals a market that is bifurcating. The model layer is becoming commoditized. The value is migrating to the application layer and the infrastructure layer. This is a familiar pattern. It happened with databases. It happened with cloud computing. It is happening with AI.
The Commoditization of the Model Layer
Open-source models are becoming the default for a growing range of tasks. This is not a niche. It is a mainstream shift. The implications for closed-source providers are profound. They cannot compete on price. They must compete on value. This means focusing on high-complexity tasks, enterprise-grade security, and reliability. The premium must be justified by capability.
Anthropic is already executing this strategy. The high expenditure share is evidence that the market is willing to pay for quality. The question is whether this premium is sustainable. If open-source models continue to improve, the premium will erode. The moat will narrow.
The Rise of the Inference Layer
The cost advantage of open-source models is not a subsidy. It is an engineering achievement. DeepSeek's MoE architecture and MLA attention are genuine innovations. This means the competitive advantage is in the inference optimization, not just the model weights. The companies that can optimize inference will win the cost war. This is a technical challenge, not a marketing challenge.
The Application Layer Opportunity
The lower cost of open-source models expands the total addressable market. Applications that were previously uneconomical are now viable. Small and medium-sized businesses can now integrate AI without prohibitive costs. This is a tailwind for the entire ecosystem. The pie is growing, even if the distribution is uneven.
The DeepSeek Phenomenon
DeepSeek's rise is the most significant data point in the Vercel dataset. It is not just a Chinese model. It is a demonstration that open-source can compete on capability, not just price. The token share growth is a vote of confidence from the developer community.
But there are caveats. The token consumption may be concentrated among Chinese developers. The international adoption may be lower than the aggregate data suggests. The model's performance on non-Chinese tasks may be less competitive. The data does not provide geographic segmentation. The "second-largest provider" status may be inflated by regional concentration.
This is a critical unknown. The global AI market is not homogeneous. Regional preferences and capabilities matter. The Vercel data is a global snapshot, but it does not tell us where the volume is coming from.
The Google Warning
Google's decline in the Vercel rankings is a warning sign. The company has world-class research capabilities. Its models are competitive on benchmarks. But the developer experience is lacking. The API pricing is not aggressive. The tooling is not integrated. The ecosystem is not sticky.
This is a lesson in execution. Research leadership is not sufficient. The market rewards products, not papers. Google's position in the AI model market is weaker than its reputation suggests. The Vercel data is a reality check.
The Bull Market Context
We are in a bull market for AI. The hype is real. The funding is flowing. The valuations are expanding. But the Vercel data is a reminder that the fundamentals are shifting. The market is not a monolith. The winners are not predetermined. The data is the ultimate arbiter.
In my experience auditing ICO whitepapers in 2017, I learned to be skeptical of narratives. The projects with the best marketing often had the worst tokenomics. The projects with the most modest claims often had the most sustainable models. The same principle applies here. The AI model market is full of narratives. The Vercel data is a check on those narratives.
The Takeaway: The Next Signal
The Vercel data is a snapshot, not a verdict. The next signal will come from the capability frontier. If open-source models close the gap on complex reasoning and tool use, the expenditure share will shift. The 91.4% figure will erode. The value concentration will disperse.
I will be watching the next model releases from DeepSeek and the open-source community. I will be tracking the token share and expenditure share on platforms like Vercel. I will be looking for the inflection point where the value gap narrows.
History repeats not by fate, but by flawed code. The AI model market is writing its own code. The question is whether the open-source community can write the next chapter. The data will tell us. It always does.
Trust is a variable, not a constant in this market. The only constant is the data. And the data is clear: the volume has migrated. The value has not. Yet. The question is not whether the value will follow. It is when. And which models will capture it.
The next quarter will be decisive. The next model release will be a signal. The next Vercel update will be a confirmation. I will be watching. The data will not lie.