Who Will Earn on the Second Wave of AI
An analytical essay on the new AI balance of power
Roman Z., VC market analyst
In the summer of 2026, the biggest news about artificial intelligence looks political. Mutual bans, export controls, and a bill to shut models down. Yet behind almost every story sits an ordinary commercial calculation. Let us work out what exactly it is.
Introduction
For several years the industry lived by simple logic. The smarter the model, the more valuable it is, and whoever builds the most capable neural network first takes the market. That logic is now failing. American labs still hold the strongest models, but Chinese competitors are rapidly increasing their share of real-world usage, and the reason lies less in quality than in the price per query. Almost the entire new agenda around AI stems from this divergence, the political one and the investment one alike.
Bans instead of competition
Over July and August, the two sides traded restrictions. Chinese regulators warned about the risks posed by Anthropic's products, and Alibaba barred its employees from wusingClaude. Officially, this concerns data security, but commerce stands behind it too. American labs fear distillation, the training of someone else’s model on their prompts and responses. Alibaba engineers had been using Claude openly for a long time, and that has now become a legal risk.
The accusations reached the level of governments. The White House stated that Kimi K3 had been trained on Claude Fable responses. Proving that is difficult, but the accusation already works as an argument in negotiations. Nothing beyond words has followed so far, because openness itself got in the way. The models ship under free licenses and are hosted on Western platforms without any trouble.
Europe chose a third path and strengthened its enforcement powers under the AI Act, which sorts AI systems by risk level and sets separate requirements for each. The EU has no AI leaders of its own, and regulating other people’s companies is easier than catching up. Beijing, Washington, and Brussels all follow the same design. Restrictions appear wherever cheap open models start competing with expensive closed ones.
Smart does not mean efficient
The market has split into an expensive premium tier and a cheap mass-market tier, with the latter growing faster.
Frontier models really are very smart, but smart does not mean profitable. On a complex task, their capabilities can justify the high price. For a simple query, such power is often excessive because the model uses more tokens, and each one costs more. And simple queries outnumber the complex ones by an order of magnitude.
The price difference is already noticeable. A million output tokens from Kimi K3 cost around $15, against $50 for Claude Fable 5.
The consequences show up best on routers, the platforms that let you work with dozens of models through a single interface and pick one for a specific task. According to data from OpenRouter, Chinese open-weight models accounted for less than 2% of weekly traffic at the end of 2024, and by April 2026, their share had surpassed 45%. Over the same stretch, the volume of requests on the platform itself grew fourfold. Open models are therefore gaining share in a market that is expanding rapidly on its own, rather than winning pieces away from stagnant demand.
One important caveat belongs here. A significant portion of the Western labs’ traffic never reaches routers at all, since they host their own models, and the premium price makes a single token cost as much as a dozen cheap ones. So these figures show the direction of movement rather than the exact balance of power. The direction is fairly simple. The market has split into an expensive premium tier and a cheap mass-market tier, with the latter growing faster.
The quality gap is narrowing
As long as the quality gap stayed noticeable, the premium segment was protected. Over the summer, that protection weakened. In the composite Artificial Analysis index, the best open model, Kimi K3 from Moonshot AI, trails the leading Claude Opus 5 by just three points out of 63. Among the 14 best models in the ranking, six are Chinese open-weight models.
Kimi K3 delivers results beyond benchmarks as well. The model works as an agent when writing code and can handle large repositories, entire project codebases where dozens of connected files must be managed at once. These are already thoroughly applied tasks.
Qwen3.8 from Alibaba, released soon after, holds seventh place and matches Kimi on programming tasks, but it burns more tokens than average and costs more than most open competitors. A useful reminder that "open" no longer automatically means "cheap".
For Anthropic, this stings in particular. The company had to double its token prices because capacity could not meet the full demand. The market accepted the increase, and Anthropic became the first of the large AI labs to reach operating profit. How sustainable that model proves to be remains an open question. The market accepted the high price at the peak of the hype, and it may well be unwilling to keep paying it over the long run, especially once open models of comparable capability appear alongside at a third of the cost.
Faster equals cheaper
Closed models hold two key advantages, and we have already covered the first. On genuinely complex tasks, a strong model can spend fewer tokens and end up cheaper. The second advantage is speed. A user cares not only about price but about how long the task takes. Getting 10,000 lines of code for $5 makes sense if it takes five minutes rather than several hours.
Speed depends on hardware, model architecture, and server software. Kimi K3 on Moonshot servers produces 40 to 80 tokens per second, while Claude and ChatGPT on their own servers produce more than 200. On a complex task, a closed model can therefore spend fewer tokens and run faster at the same time.
This advantage is not final either. On powerful infrastructure, open models close the speed gap while keeping their lower price.
On genuinely complex tasks, a strong model can spend fewer tokens and end up cheaper.
The money is moving into infrastructure
The fight for inference efficiency (a trained model at work when it answers a query) has redrawn the hardware market. At the end of 2025, NVIDIA signed a non-exclusive licensing agreement with Groq for its inference technology and hired the founder along with the team. Formally, Groq kept its independence, but its technology and key specialists now work for a former potential competitor.
Demand for fast, cheap inference is lifting Cerebras and Etched, while Anthropic and OpenAI have already begun developing their own chips. In August, Anthropic officially announced the creation of a team to develop custom hardware, and OpenAI presented its own inference chip, Jalapeño.
This, to my mind, is where the main conclusion of the whole story lies. The first wave of AI investment was extremely concentrated, and the money went above all into the labs themselves. The second wave looks far broader because the focus now encompasses physical AI, developers of optimized chips, data center infrastructure, and the energy that powers those data centers
Crusoe and Lambda illustrate this well. The first works on the energy side of the question, from siting data centers next to cheap power sources to generating its own power with turbines. The second rents out compute for training and inference. Both grew on the infrastructure layer rather than on their own models.
In parallel, the question of sovereignty is growing sharper. Compute is scarce; a significant share of it is held by a handful of American labs, and those same labs largely set the price of access to models for businesses and entire countries. No surprise that talk of national compute capacity has moved from presentations into budgets.
The boom has its side effects too. Salaries in the AI sector have driven up inflation in California, and if Anthropic and OpenAI really do go public, their combined market capitalization will exceed the sum of all venture exits over the past 25 years, creating a high risk of capital concentration.
The smartest one does not win
The competition between cheap and expensive models looks like geopolitics, though it began as an economic issue.
At the start of this text, I said that a commercial calculation stands behind the political news. Here it is. The race for model intelligence has run into the economics of using those models. A smart model pays off on complex tasks and loses on simple ones, which outnumber them by an order of magnitude.
One question remains, namely why the dispute fell so neatly along the line between China and the US. Here cause and effect have swapped places. China bets on open models not out of ideology, but because it has spent several years under American restrictions that closed the usual route to the Western market for its developers. A free license became one of the few ways to win back that market, and a strategy adopted out of necessity proved to be a winning one. That is why the competition between cheap and expensive models looks like geopolitics, though it began as an economic issue.
It is too early to bury closed models. Anthropic answers the pressure with more than prices. The company is developing its own chip and is already hiring a chip design team. If the cost and speed of a token have become the main competitive field, the problem cannot be solved in model architecture alone. Labs that close the chain from chip to product on themselves will gain a margin of safety that those renting someone else’s capacity lack.
For an investor, the conclusion leans optimistic. The fight now covers not only intelligence but the cost and speed of a single token. Labs are not the only ones optimizing those. Chip makers, data centers, and energy companies do it too. That is exactly where I would look for the next layer of winners.
Қазақша