Why Wall Street Is Bullish on AI Chips After Kimi K3 Release

Source Tradingkey

TradingKey - Chinese AI large language models have once again become the focus of global market attention.

Recently, Chinese AI developer Moonshot AI officially released its latest-generation large model, Kimi K3. With its open-weights architecture, performance close to that of the world's leading models, and more competitive inference costs, Kimi K3 quickly sparked heated discussions in overseas tech circles and capital markets.

Following the announcement, the market briefly worried whether the logic of global tech giants continuing to invest hundreds of billions of dollars to build AI infrastructure would be impacted as more high-performance, low-cost AI models emerge, and the US AI chip sector also experienced a noticeable correction.

However, an increasing number of analysts believe that Kimi K3 may not necessarily be a negative for the AI chip industry. Instead, improved model performance and lower usage costs could lower the barrier to AI applications, bringing more users, higher call frequencies, and larger-scale inference demand. The "Jevons paradox" in economics may be playing out once again in the AI industry.

Why Kimi K3 Is Drawing Global Market Attention?

Kimi K3 is the flagship model launched by Moonshot AI in July 2026, primarily targeting application scenarios such as long-cycle programming, complex knowledge work, and AI agents.

According to disclosures by Moonshot AI, Kimi K3 adopts a Mixture-of-Experts (MoE) architecture, featuring 2.8 trillion total parameters and a 1-million-token context window, and natively supports visual understanding. Among its 896 expert modules, only 16 are activated per calculation to improve model capacity and inference efficiency. It is worth noting that 2.8 trillion is the total parameter scale, which does not mean that every generated token invokes all parameters.

Kimi K3 has performed outstandingly in frontend programming evaluations, once ranking near the top of the Frontend Code Arena leaderboard with a score of 1,679. However, this achievement is confined to specific programming tests and does not equate to comprehensively outperforming the flagship models of OpenAI and Anthropic in terms of overall capabilities.

Pricing is also a key reason for the attention surrounding Kimi K3. Official API pricing shows a rate of $3 per million input tokens—dropping to $0.30 upon a cache hit—and $15 per million output tokens. For AI agents that need to repeatedly retrieve codebases, enterprise documents, and long contexts, the caching mechanism can significantly reduce input costs.

However, a 'lower unit price per token' does not equate to a proportional decline in the actual cost of all tasks. Model output length, inference time, cache hit rate, and task completion rate will all impact the final cost.

Therefore, the real change brought by Kimi K3 is not simply launching a price war, but rather driving the broader adoption of advanced models with a higher price-to-performance ratio.

Why Low-Cost Kimi K3 May Increase Chip Demand?

After the release of Kimi K3, some investors worried that improved model efficiency would reduce GPU usage. However, judging from actual operations, a decline in unit inference costs does not directly equate to a drop in total chip demand.

Kimi K3 adopts a sparse Mixture of Experts (MoE) architecture, activating only a subset of expert modules during each inference, which helps control the computational overhead per run. However, the model still needs to store and schedule massive weight data, while processing contexts of up to 1 million tokens. In large-scale deployments, such tasks still rely on GPUs, AI accelerators, high-bandwidth memory, high-speed networks, and server clusters.

More importantly, as model costs decrease, the number of users and frequency of use may grow even faster. Following the launch of Kimi K3, user request volume rapidly approached the capacity limit of the existing computing clusters. Moonshot AI temporarily suspended some new user subscriptions to prioritize computing power resources for existing users. This indicates that as models become cheaper and easier to use, new demand can quickly consume the computing power saved by efficiency gains.

This is precisely the core of the 'Jevons Paradox': as the efficiency of resource utilization improves, the unit cost falls, but because the scope of application expands, its total consumption may actually increase instead.

In the AI industry, improvements in model efficiency can lower the barrier for enterprises to deploy intelligent customer service, coding assistants, office agents, financial analysis, and medical assistance systems. As more applications run continuously, the daily volume of inference requests will surge, ultimately driving continued growth in demand for computing power, memory, and data center infrastructure.

Low-Cost Models Do Not Mean Computing Power Investment Is Without Risk

While the Jevons paradox provides an optimistic explanation for AI chip demand, it is not an inevitable outcome. If the rate of model efficiency gains outpaces user and application growth in the long run, the decline in computing power required per task could still compress a portion of hardware demand.

AI infrastructure investment must ultimately face the test of commercial returns. If companies pour massive amounts of capital into building data centers but fail to generate sufficient revenue through subscriptions, advertising, cloud services, or enterprise software, the growth rate of capital expenditures may slow. The extent to which different chip companies benefit will also diverge, and not all AI-concept stocks will grow in tandem.

Kimi K3 could also intensify price competition in the model market. As open-source or open-weight models increasingly close the gap with the capabilities of closed-source models, the API profit margins of AI companies may come under pressure. To lower costs, model developers will place greater emphasis on custom chips, quantization technologies, inference optimization, and multi-vendor strategies, which could weaken the bargaining power of general-purpose GPU makers in certain markets.

Therefore, the impact of low-cost models on the AI chip industry is not a one-way positive. While it expands the AI application market on one hand, it also forces infrastructure companies to improve energy efficiency and reduce unit computing costs on the other.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Bitcoin Security Consortium launches with $15M support for quantum, long-term securityNine of the largest firms in institutional Bitcoin have announced the roll-out of the Bitcoin Security Consortium. The announcement came on Thursday, July 23.  The nine firms committed a total of $15 million over three years to help developers and researchers keep the network safe, with emphasis on post-quantum cryptography.  The group includes companies with...
Author  Cryptopolitan
20 hours ago
Nine of the largest firms in institutional Bitcoin have announced the roll-out of the Bitcoin Security Consortium. The announcement came on Thursday, July 23.  The nine firms committed a total of $15 million over three years to help developers and researchers keep the network safe, with emphasis on post-quantum cryptography.  The group includes companies with...
placeholder
AMD launches new AI server in direct challenge to Nvidia's AI dominanceAdvanced Micro Devices (AMD) CEO Lisa Su told a San Francisco audience on Thursday that AMD’s Helios rack scale system is in full production, setting the chipmaker up to assess the AI data center business that Nvidia has controlled almost by itself. It is the first time AMD has fielded a complete server cabinet built...
Author  Cryptopolitan
20 hours ago
Advanced Micro Devices (AMD) CEO Lisa Su told a San Francisco audience on Thursday that AMD’s Helios rack scale system is in full production, setting the chipmaker up to assess the AI data center business that Nvidia has controlled almost by itself. It is the first time AMD has fielded a complete server cabinet built...
placeholder
Bullish XRP Chart Clashes With an ETF Warning, Who Wins?XRP (XRP) price is holding just above $1.13 after a mild pullback, keeping a bullish chart structure alive even as institutional demand shows signs of cooling.The token has slipped since July 21, yet
Author  Beincrypto
20 hours ago
XRP (XRP) price is holding just above $1.13 after a mild pullback, keeping a bullish chart structure alive even as institutional demand shows signs of cooling.The token has slipped since July 21, yet
placeholder
Google Is Up $94 Billion on SpaceX But Not for the Reason You ThinkGoogle just revealed it holds about $94 billion of SpaceX stock. The win came from one bet it made back in 2015.That sounds like a giant new investment, but it is not. Google made this bet more than t
Author  Beincrypto
20 hours ago
Google just revealed it holds about $94 billion of SpaceX stock. The win came from one bet it made back in 2015.That sounds like a giant new investment, but it is not. Google made this bet more than t
placeholder
SpaceX is a Warning For Crypto and Tech Stocks, Peter Schiff SaysSpaceX stock (SPCX) closed just above $115 on Wednesday, nearly 20% below its June IPO price, while its 2056 bonds sank to a record low below 89.The dual decline pushed the bonds’ yield to worst to 7.
Author  Beincrypto
20 hours ago
SpaceX stock (SPCX) closed just above $115 on Wednesday, nearly 20% below its June IPO price, while its 2056 bonds sank to a record low below 89.The dual decline pushed the bonds’ yield to worst to 7.
goTop
quote