Thu 06 Aug 2026 / 09:45 ET
Kernel
Long Reads 3 min read

Kimi K3 distillation debate puts US AI terms of service under pressure

Ben Thompson proposes fair-use protections and limits on anti-distillation terms as Kimi K3 renews anxiety over Chinese AI models.

Theo Lindgren

By Theo Lindgren / Columnist

Ben Thompson has turned the Kimi K3 distillation fight into a policy argument: American open-weight model developers are bound by frontier labs’ API rules, while Chinese competitors can treat those same systems as training material and ship usable alternatives faster.

Writing at Stratechery, Thompson argued that US open-weight model makers end up at a disadvantage because they must obey terms of service from leading AI labs. In his view, those restrictions make their models weaker than Chinese alternatives and push them into an absurd loop: instead of learning directly from US frontier models, they learn from Chinese models that may already have learned from those US systems.

What is model distillation?

Model distillation is a training shortcut in which a developer queries a stronger model, captures its answers, and uses those answers to train another model. For readers who want the machinery underneath that process, Kernel has a guide to how LLMs work.

Thompson’s complaint is not that Chinese models such as K3 are mere copies. Daring Fireball, discussing Thompson’s piece, also stressed that leading Chinese models are not strong only because of distillation. The argument is narrower: distillation appears to be a meaningful part of how those models keep arriving within roughly six to nine months of the US frontier, according to Daring Fireball’s account.

The asymmetry is the problem. Thompson said Chinese developers effectively treat available models as open regardless of contractual limits. Developers in the US and other countries that take copyright and terms of service seriously do not have the same freedom, so they may have to wait for Chinese releases such as K3 before building competitive open-weight models of their own.

What policy change did Thompson propose?

Thompson called for a US law with two pieces. First, it would state that gathering data to train models is fair use. Second, it would bar terms of service that prohibit distillation, at least for US companies.

His rationale is that stopping distillation is technically hard because the act can look like ordinary API use: a system sends prompts and receives responses. If frontier labs can train large language models on material from the open internet, Thompson argued, downstream model builders should not be blocked from using what those frontier systems produce as fuel for further development.

That position would cut against the interests of companies such as OpenAI and Anthropic, which Daring Fireball described as objecting when their frontier models are distilled against their rules. The post was openly unsympathetic to those objections, pointing to the tension between labs that learned from broad public data and then try to contractually fence off what others can learn from their models.

The fight is partly about copyright, partly about contracts, and partly about industrial policy. Thompson’s concern is that strict private terms may leave US open-weight developers dependent on Chinese releases, even when the underlying technical advantage began in the United States.

This story draws on original reporting from Daring Fireball.

More Long Reads/

view all ↗