As Kimi K3 breaks through the closed source premium, open weight AI is detonating the “model routing revolution”! As tokens get cheaper, demand for AI computing power explodes

Zhitongcaijing · 2d ago

The Zhitong Finance App learned that the recent AI big model technology evolution path and market penetration trend can be described as shocking to institutional investors focusing on the valuation prospects of closed source AI application development leaders such as OpenAI and Anthropic, especially the world, including Nvidia (NVDA.US), Microsoft (MSFT.US), Facebook parent company Meta (META.US), established US tech giant IBM (IBM.US), and Silicon Valley venture capital giant Andreessen Horowitz Leading forces in the technology industry are urging the US government to support the development trajectory of open-weight AI (open-weight AI) AI big model technology. More than 20 of the world's top technology companies, including Nvidia and Microsoft, signed an open letter supporting the open weighted AI model. OpenAI and Anthropic did not sign the letter, but Musk, the world's richest man and founder of SpaceX, and those at the helm of tech giants such as Microsoft CEO all publicly expressed their support.

According to reports, in the AI developer ecosystem, open-weight AI usually refers to an “open weight AI model”, that is, developers can download parameters after model training and deploy, fine tune, and reason in their own server or cloud environment; however, model developers do not necessarily disclose training data, data processing methods, training code, and complete architecture details, and may also restrict commercial use or specific applications through licenses.

Open source AI in the strict sense of the word is broader: in addition to weight, it should also provide information sufficient to allow users to research, modify, and reproduce the system's code, architecture, and training data, and guarantee basic freedom of use, research, modification, and redistribution. Therefore, open weight is “the finished model can be used and modified”, while true open source AI further opens up “how the model is manufactured”; the former can be part of the latter, but open weight is not necessarily equal to open source all technical aspects and code.

According to information, this important AI technology industry open letter released last Friday comes at a time when the White House is considering whether to ban China from opening up the weighting model due to national security concerns. More than 20 companies signed the letter, including leaders in the field of AI application software such as A16z, Dell, Microsoft, Meta, Nvidia, and Palantir.

For the AI computing power industry chain and AI super bull market investment trend that global investors are focusing on, opening up weight, model routing, and improving architectural efficiency will reduce the cost of intelligence per unit, but may amplify total computing power demand through the Jevins paradox. As simple call costs decrease, enterprises will deploy more continuously running agents, parallel sub-agents, long context analysis, code automation, and real-time multi-modal services; a single task may consume less computing power, but the number of tasks, inference steps, and deployment nodes may grow faster. Even if only a small number of experts are activated, very large MoE models such as KIMI K3 still need to distribute huge weights in high-capacity memory and high-speed interconnect clusters. Therefore, the popularity of open models will spread computing power requirements from a few closed source laboratories to cloud service providers, sovereign clouds, giant AI data centers for enterprises, and local inference clusters.

Recently discussed, the so-called Jevons Effect (Jevons Effect), also known as the Jevons Paradox, is a counterintuitive economic theory: when current technological advances improve the efficiency of the use of certain resources (such as energy, raw materials, or AI computing power infrastructure resources), it will reduce the unit cost, thereby stimulating a large-scale expansion of market demand, which ultimately causes the total consumption of all types of resources to increase rather than decrease. The concept was first proposed by the English economist William Stanley Jevons (William Stanley Jevons) in the book “The Coal Problem” in 1865.

Why are US tech giants suddenly supporting open weighted AI?

Microsoft and other co-signers stated in this open letter: “The open weighted AI model has expanded opportunities for global companies to participate in economic prosperity in the AI era.” The letter stated that various organizations can develop on the basis of advanced AI models without training the model from scratch or paying the expensive token price for cutting-edge models.

The tech giant added that open weights can promote competition and ensure that the benefits brought by AI are shared more widely rather than concentrated in the hands of a few players.

“Open weighting allows every organization to match the right model for the right task at the right cost, leaving the ability at the cutting edge of scale to real cutting edge problems, while integrating, operating, and deploying efficient and specialized models in all other scenarios.”

According to information, Nvidia CEO Hwang In-hoon has always been an active proponent of open weighted AI. While sharing the joint letter, he himself wrote, “I am sharing a letter from Nvidia participating in the joint signing, which explains why it is important to open up the big AI model. AI will ultimately transform every industry, power every enterprise, and have every country participate in and lead construction.”

Unlike closed source AI systems such as those led by OpenAI and Anthropic, the open weighting model allows developers and enterprises to download almost all trained parameters — the so-called “big model weights” — and run on AI computing power infrastructure created exclusively for their own systems. Essentially, the idea is similar to the open source software movement that changed the global computer industry decades ago.

Open weighting models are different from fully open source AI. Open weight models usually disclose model weights after training is completed, and everyone can download and complete one-click deployment for free, but they do not necessarily disclose the underlying training data or all source code for large AI models at the same time.

Training advanced AI models can cost billions of dollars. The open letter argues that an open weighting model can largely solve this problem, because developers can develop or deeply customize training based on existing AI models without having to build a model from scratch. For smaller companies, this has drastically lowered the threshold for participating in product contests in the AI era.

According to the letter, the open weight model can not only promote competition among AI developers, but also promote healthy and active competition among cloud computing service providers, chip companies, software providers, and AI application developers.

“This kind of competition can stimulate innovation, reduce costs, and make the benefits brought by AI widely benefit the entire economy,” because open weighting AI drastically reduces a series of costs for developing AI products.

In contrast, developers of closed AI models such as OpenAI, Anthropic, and Alphabet's Google (GOOGL) .US clearly have strong economic incentives to continue to maintain exclusive control over their top technology.

So how do these tech giants, including Nvidia, view the security risks associated with open weighted AI models? The open letter added that the cybersecurity defense field requires priority access to these most advanced open source or closed source AI capabilities in order to cope with increasingly complex cybersecurity attacks in the AI era. The researchers in the letter believe that the big open weight AI model can allow more researchers to identify bugs, improve security protection systems, and conduct large-scale and independent tests on the system without completely relying on the original model developers.

Kimi set off the “open weight+low price token” shock wave, and model routing takes over the AI entrance

The real impact of Kimi is not only “lower price”, but the high performance cost ratio, open weight, long context, and development interface compatibility all lower the model migration threshold. The official price of Kimi K2.6 is about $0.95 and $4 per million input and output tokens, respectively, and the cache input is only $0.16; Kimi K3 provides the context of 1 million tokens, and the prices on OpenRouter are $3 and $15. As a result, “low price” is not the absolute lowest of all models compared to its level of capability; more importantly, open weights allow enterprises to deploy on their own infrastructure, regional clouds, or third-party inference platforms, reducing their reliance on a single API vendor. Kimi K2.5 has accumulated more than 12.3 billion token usage on OpenRouter, indicating that China's open model is rapidly entering the global AI developer toolchain and is no longer just a local AI chatbot window product.

The so-called “model routing” means that in the future, enterprises will no longer bind all tasks to OpenAI, Anthropic, or a specific model, but will dynamically select models based on quality, cost, delay, compliance, context length, and usability when the request arrives: simple classification, summary, and customer service will be handed over to the low-cost open model. Complex reasoning, key code, and high-risk decisions will only invoke expensive cutting-edge models.

image.png

The Model Router created by Microsoft already supports the three routing modes of cost, quality, and balance, and provides automatic failover; Microsoft believes that the router can direct 60% to 80% of traffic to cheaper models without measurable loss in quality. AWS internal testing showed that intelligent routing can save about 16% to 56% of costs in different model families, and the savings reached 63.6% in some RAG tests. This means that the core entry point for AI applications is moving from a “model brand” to a routing and orchestration control layer, which controls traffic allocation, model negotiation, and actual token consumption.

This is where the “Jevins Paradox” is likely to repeat itself in the AI industry: model routing, caching, quantification, and open weighting reduce the marginal cost of each inference, but cheap intelligence will spawn new requirements that were previously economically unfeasible — continuously running enterprise intelligence, hundreds of internal inference and self-checks, parallel sub-agents, analysis of millions of token documents, real-time voice and video understanding, and back-office AI embedded in every software operation. As long as the flexibility of token demand on the price is greater than 1, that is, after the unit price drops by 50%, the total computing power consumption and total inference expenses will increase.

Therefore, the low-cost model may raise concerns that “computing power demand will peak” in the short term, but in the long run, it is more likely to spread AI from a few high-value tasks to billions of daily workflows, transforming phased capital expenditure in the training era into continuous, distributed, and high-frequency computing power consumption in the inference era. The Open Weights Joint Letter itself is also the key to economic sustainability, defined as allowing cutting-edge capabilities to address only cutting-edge issues, while allowing efficient dedicated models to cover a large number of everyday tasks.

AI chips, storage, optical interconnects, AI control planes, and vertical applications

For the AI computing power industry chain, this does not mean that the demand for AI computing power has disappeared; rather, the computing power structure is shifting from centralized training to training and extensive reasoning. Frontier training still requires high-end GPUs, advanced manufacturing processes, HBM/enterprise-grade DRAM/data center NAND, advanced packaging, and high-speed optical interconnection; open models and model routing will expand the inference deployment of enterprise private clouds, sovereign AI, regional clouds, and local data centers, increasing the demand for inference GPUs, dedicated accelerators, server CPUs, high-end DRAM/NAND storage components, and optical interconnect systems such as Ethernet infrastructure and optical modules, and power and liquid cooling systems.

Furthermore, technical pricing anchors will also shift from simply pursuing peak FLOPS to tokens per watt, tokens per dollar, memory bandwidth, KV Cache efficiency, MoE sparse activation, and cluster utilization. It should be emphasized that open weighting does not mean that it can run at a low cost on a PC: Kimi K3 has 2.8 trillion total parameters and 1 million token contexts, and its full deployment is still a large-scale rack or data center level project, which in turn is beneficial to suppliers with large-scale reasoning, cluster scheduling, and energy infrastructure capabilities.

OpenAI and Anthropic will not lose all room for growth as a result, but the token toll gate model where “all requests call the closed source model at a high price” will be structurally eroded. Anthropic's Claude Opus 5 is priced at $5 and $25 per million input and output tokens, and Sonnet 5 is being promoted at $2 and $10; in the face of open models and automatic routing such as Kimi, Closed Source Labs must prove that its cutting-edge capabilities, reliability, security compliance, tool ecology, and end-to-end intelligence can create business value significantly above the price difference.

According to Goldman Sachs, the AI computing power super bull market is far from over. Instead, it has moved from the “AI chip purchase frenzy” to the second stage of “large-scale AI factory construction” — that is, the next round of excess alpha revenue will no longer only belong to the list of the strongest leaders in the AI GPU/AI ASIC field, but will also spread systematically to data center high-performance CPUs, DRAM/NAND/HBM storage, AI PCBs, liquid cooling systems, data center optical interconnection systems, ABF carriers/glass substrates, MLCC, electronic distribution, and extensive foundry of “AI factories” full-stack AI computing Power infrastructure layer.

According to a recent research report by senior analyst Brian Nowak from Wall Street financial giant Morgan Stanley, leading the analysis team, the 2027/2028 capital expenditure forecasts for the five largest hyperscale cloud computing and vendors (Meta, Amazon, Microsoft, Google, and SpaceX) in the global market (Meta, Amazon, Microsoft, Google, SpaceX) have been significantly raised again, reaching approximately $1.2 trillion and $1.4 trillion, respectively. The agency's capital expenditure forecast for major US tech giants in 2026 was drastically raised from 433 billion US dollars a year ago to 805 billion US dollars.

image.png

Morgan Stanley's latest study raised Meta's 2027 and 2028 capital expenditure forecasts by 29% and 22%, respectively, to US$225 billion and US$250 billion; Amazon's corresponding forecasts were raised 15% and 29% to US$308 billion and US$318 billion. Morgan Stanley said that the capital expenditure supercycle is not over yet, but 2026 and 2027 are probably the steepest years of growth. After 2028, what determines the stock price will no longer be just “who spends the most money,” but “who can quickly turn AI computing power resources into revenue, profit, and free cash flow.”

Although MoE, model routing, quantization, and caching will reduce the cost per token, they may spawn more resident agents, longer inference chains, and higher call frequencies through the “Jevens Paradox”, causing total computing power and total electricity demand to continue to rise.

Another Wall Street financial giant, Citigroup's latest research report shows that the arms race in the AI era is shifting from “who has the smartest model” to “who can continuously produce intelligence at the lowest cost and with the highest efficiency under physical constraints”: open weight models such as Kimi K3 are rapidly approaching the frontier of closed source, which means that model capabilities accelerate commercialization, but parameter scale, long context, and multi-step agent reasoning expand simultaneously, causing bottlenecks to move from simple flops to HBM capacity and bandwidth, high-speed GPU interconnection, cluster scheduling, and power access. Nvidia research also points out that when model size, sequence length, and batch expansion, HBM often becomes the main expansion constraint; IEA predictions show that AI data center electricity consumption is growing significantly faster than overall electricity demand, while the power grid construction cycle is generally longer than the data center deployment cycle.

If the AI superbull market continues, the future investment trend in the AI computing power industry chain may focus on AI chips, storage, optical interconnection, AI control planes, and vertical applications for a long time. This is why Wall Street financial giants such as Morgan Stanley, Bank of America, and Nomura continue to be optimistic about leaders in the AI computing power industry chain such as Nvidia, Micron, SK Hynix, and Intel. In other words, computing power infrastructure leaders who master AI chips, storage components, network infrastructure, and energy bottlenecks; AI control planes that master routing, evaluation, inference optimization, security and observability; and vertical applications that can use token deflation to expand usage and form real cash flow. What needs to be wary of is software companies that lack proprietary data, workflow barriers, and distribution channels, and are only built on a single closed source API — the more mature the model route, the easier it is for its products to be replaced, and the easier it is for profit margins to be swallowed up by underlying price wars.