Moonshot AI introduced Kimi K3 on July 17 as fresh “evidence” that Chinese developers could approach the AI frontier at lower advertised prices. Within days, demand pushed its GPU clusters close to capacity, forcing it to pause new consumer subscriptions.
Something very similar had happened when the Chinese DeepSeek V3 was released. It had suffered a similar capacity crunch during its viral rise in early 2025.
The episode has again reminded investors of the split in AI economics. Competition is making models harder to sell at premium prices, but compute remains scarce. Nvidia Corporation (NASDAQ:NVDA) and Microsoft Corporation (NASDAQ:MSFT) may be better positioned for that world than the labs financing the model race.
Kimi Is Not as Cheap as It Looks Against All Frontier Models
K3’s API costs $3 per million input tokens and $15 per million output tokens. Artificial Analysis nevertheless estimates an average cost of $0.94 per Intelligence Index task. That compares with $0.55 for GPT-5.6 Terra and $1.04 for GPT-5.6 Sol, both for maximum reasoning effort, so not exactly as cheap as it’s hyped up to be against frontier OpenAI models. That said, K3 is substantially cheaper than Claude Fable 5, which is at $2.75.
The gap with GPT-5.6 is partly a token-efficiency problem. On the same standardized workload, K3 generated roughly 25,000 output tokens per task, including about 18,000 internal reasoning tokens, while GPT-5.6 Sol used 15,000 output tokens in total. K3 therefore burned about 67% more output tokens. Since reasoning tokens are billed at the output rate, K3’s 50% per-token output discount to Sol narrowed to about 17% for the output portion of that workload.
LocalLLaMA users on Reddit also noted that a 2.8-trillion-parameter model would be difficult for most people to run locally, a view that was echoed in the sub after GLM 5.2’s release as well. The cost of hardware has a huge discrepancy with the pace at which locally runnable models are advancing.
The threat to frontier U.S. labs, though, is not as much as K3’s current efficiency as it is the speed with which competing models are approaching the frontier.
That is the core of investor Gavin Baker’s argument. K3 does not need to be the cheapest model in the field to hurt frontier intelligence economics. A near-frontier model promised as open weights can limit what closed labs charge, compressing model-level margins. As AI becomes cheaper to deploy, adoption and inference demand can increase, and more value could shift toward semiconductors, power, cloud infrastructure, and software. K3’s current compute inefficiency can therefore be negative for model-lab margins while remaining positive for infrastructure demand.
Still, frontier developers are unlikely to respond by reducing investment. Their products lose relative value whenever a competitor advances, creating pressure to train another generation. Lower inference prices can also push for more applications and coding agents that make repeated model calls.
OpenAI and Nvidia have announced plans for at least 10 gigawatts of AI systems, while Anthropic has signed agreements for up to 10 gigawatts through Amazon, Google and Broadcom. That capacity will train future models and serve agentic demand. OpenAI says median Codex token use among its researchers rose 56-fold in seven months.
Cheaper Models Still Need Expensive Hardware
SemiAnalysis says that K3 is positive for Nvidia Corporation (NASDAQ:NVDA) despite its more efficient attention mechanism. The model’s 2.8 trillion parameters occupy more than 1.5 terabytes of HBM capacity and create heavy requirements for GPUs, storage, and high-speed networking. Moonshot’s capacity shortage supplied an immediate demonstration of that inference demand.
Nvidia Corporation (NASDAQ:NVDA) is positioned across the stack through its accelerators, NVLink interconnects, networking products and CUDA software.
Microsoft Can Change the Model Without Losing the Customer
Microsoft (NASDAQ:MSFT) faces some exposure to weaker model economics through its OpenAI relationship. Its larger advantage is controlling Azure, Foundry and the enterprise applications through which businesses consume AI.
Microsoft (NASDAQ:MSFT) Foundry can route workloads among OpenAI, Microsoft and open models based on cost and performance. Microsoft has also reportedly started using its own MAI models for selected workloads previously handled by third-party systems. If model prices fall while usage grows, Microsoft can still monetize the underlying compute and distribution.
According to Insider Monkey’s database, 282 hedge funds held Microsoft shares at the end of the first quarter of 2026, down from 312 at the end of Q4 2025.
To sum it up, Kimi K3’s release shows that the companies building the models may capture less of their profits than the businesses supplying and distributing the compute.
While we acknowledge the risk and potential of MSFT as an investment, our conviction lies in the belief that some other AI stocks hold greater promise for delivering higher returns and doing so within a shorter time frame. If you are looking for an AI stock that is more promising than MSFT and that has 10,000% upside potential, check out our report about the cheapest AI stock.
READ NEXT: 33 Stocks That Should Double in 3 Years and Cathie Wood 2026 Portfolio: 10 Best Stocks to Buy.
Disclosure: None. Follow Insider Monkey on Google News.
