Moonshot AI has put a new problem in front of OpenAI, Anthropic and the companies paying their invoices: a frontier-class model that appears much cheaper to call, while being large enough to make self-hosting a hardware project rather than a download-and-pray exercise.
The Chinese company’s Kimi K3 is an open-weight model with 2.8 trillion parameters, making it the largest open-weight model released so far, according to Tom’s Hardware. Moonshot has not yet released the weights. That is scheduled for July 27, so current performance claims depend on API access rather than independent testing of the model files.
Moonshot’s internal benchmarks put Kimi K3 in the same bracket as models including GPT 5.5 and Claude Opus 4.8, according to Tom’s Hardware. Arena.ai ranked Kimi K3 first in its Frontend Code Arena, ahead of Claude Fable 5, and said the model jumped from 18th place in the prior Kimi-k2.6 showing to first place. The same rankings do not show Kimi K3 winning across the board, and reports cited by Tom’s Hardware say it is slower than frontier models from Anthropic and OpenAI.
The price is the provocation
OpenRouter lists Kimi K3 pricing at $3 per million input tokens and $15 per million output tokens. By the same table, OpenAI’s GPT 5.6 Sol costs $5 and $30, while Anthropic’s Claude Fable 5 costs $10 and $50.
That gap is why developers and companies with large AI bills are paying attention. Tom’s Hardware reported that some organizations have already been cutting back AI use after questioning whether heavier model spending was producing better products. Uber, for example, limited developer AI use, according to the same reporting.
Kimi K3 also arrives in the shadow of DeepSeek’s R1 release in 2025, which raised similar questions about whether U.S. frontier labs were charging a premium because their systems were better, because their costs were higher, or because customers had few credible alternatives. Tom’s Hardware notes that Kimi K3 is different in a key way: DeepSeek R1 was associated with leaner training economics, while Kimi K3 is enormous.
Bloomberg reported that Kimi K3 has an unusually high sparsity ratio. In plain English, the model contains trillions of parameters, but activates only a subset of them for a given task. That can reduce compute per query. It does not remove the need to keep the whole model available in memory.
Open weights do not make it cheap to host
Bloomberg estimates Kimi K3 needs close to 1.5 TB of memory. Tom’s Hardware puts the figure at up to 1.4 TB. Either way, this is not a model most companies will run on a few spare GPUs under someone’s desk.
Deploying Kimi K3 at scale would require high-end accelerators, large memory capacity, bandwidth and interconnects. That limits the practical audience for self-hosting, even if open weights let companies avoid Moonshot’s cloud service and build their own deployment.
The hardware angle also explains why memory suppliers are watching. Tom’s Hardware reported that Chinese DRAM maker CXMT is increasing capacity and could pass Micron’s DRAM wafer capacity by the end of the year. If large sparse models become more common, demand for memory may remain strong even when per-query compute becomes more efficient.
The political reaction is already forming. Tom’s Hardware reported that Microsoft is considering Kimi K3 for Copilot, while the White House may move to restrict Chinese AI models on cybersecurity grounds. Downloadable weights would make any clean ban harder to enforce once the files are public.
Kimi K3’s early pitch is therefore awkward for nearly everyone: cheaper API access than U.S. closed models, public weights on the way, strong coding benchmark results, and hardware needs large enough to keep infrastructure vendors smiling. The model may pressure prices before it changes who can realistically run frontier AI at scale.
This story draws on original reporting from Tom's Hardware.