White House science adviser Michael Kratsios said the Chinese company Moonshot, which developed the Kimi K3, the largest open-weight LLM available, built the model by copying Anthropic’s Fable LLM while using chips that are not allowed to be exported to China.
Amid reports of debate over a ban on China’s promiscuous models that are disrupting the AI field, Kratsios wrote that “massive, covert industrial distillation aimed at stealing America’s proprietary technology and undermining American research is unacceptable.” Moonshot did not respond to questions about the training process, and Kratsios did not provide details about the source of the allegations.
Kratsios’ tweet echoed Treasury Secretary Scott Bessent’s comments that “many of our Chinese models have been found to have the watermarks of America’s large-scale language model, and that is unacceptable.” It is not clear what these watermarks consist of, and the Treasury Department did not respond to inquiries.
However, experts are skeptical that distillation, the process of querying an LLM to determine its inner workings and copying its functionality, is responsible for the advanced functionality exhibited by Kimi K3.
“I don’t think we’ll ever get a model this strong and this quickly following Fable, which does rigorous distillation,” Braden Hancock, a researcher at the Lord Institute and co-founder of Snorkel AI, told TechCrunch. “Frankly, you don’t have time. Fable just became generally available on July 1st. There’s no way we can extract this much data, train a model, and release it in two weeks.”
Nathan Lambert, an artificial intelligence researcher at the Allen Institute for Artificial Intelligence, said in a podcast published yesterday: “I’m of the opinion that the influence of distillation is becoming less and less pronounced over time as the Chinese model approaches the frontier and training regimes move toward (reinforcement learning).” “If that were the case, anyone could easily catch up to GLM and K3 by using data in distillation. But supervised tweaking alone has not and will not achieve this.”
To perform distillation, the target model must be systematically queried in the lab to generate data that can be used after training. In some cases, this may explicitly include asking the model to articulate its chain of thought in order to understand how it solves the problem. You may also use prompts and responses from your model to train new models in a process called supervised fine-tuning (SFT).
This fine-tuning process could result in a model created by a third party ostensibly claiming to be Claude. In Lambert’s view, fine-tuning is where “the model gets its manners.”
But Lambert says that as models become more complex, the benefits of SFT become less important. Extracting fable-like features may require reinforcement learning techniques. Often this means having the larger model’s agent evaluate the smaller model’s response and adjusting based on that evaluation.
More advanced technology also requires more critical infrastructure. Running reinforcement learning at scale can require tens of millions of agents. Using Frontier Labs’ API to do this “could be very expensive and probably a time bottleneck, because these models are very slow and frankly you might not even get a performance improvement.”
It looks like a previous Frontier model may have served you well. Earlier this year, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically extracting its models. Anthropic said it discovered millions of interactions between its models and users identified by these companies through IP addresses and other metadata. These queries “deviate from normal usage patterns and reflect intentional feature extraction rather than legitimate use.” Anthropic did not respond to TechCrunch’s questions about Fable’s distillation.
However, distillation is not limited to China and is considered common among AI companies. Elon Musk testified earlier this year that his company SpaceXAI extracts OpenAI models to develop Grok, and that the practice is common in the industry. For example, the line between distillation and the development of synthetic datasets can become quite blurry.
“In general, Americans underestimate the technical expertise of the Chinese team,” Hancock said. “One of the founders of Moonshot was a PhD student at CMU. They’re legitimate researchers and engineers doing solid work. … If the American model stops, I think China’s progress will slow down, but it’s still going to continue. They’re not just sitting on the tail here.”
It’s also difficult to disentangle the second part of Kratsios’ comment, that Moonshot acquired an advanced Nvidia chip, the Grace Blackwell 300, and also had access to a GB300-powered server in Thailand. Although exports of these chips to China are prohibited, a black market exists, said Sam Bresnick, a researcher at Georgetown’s Center for Security and Emerging Technologies. In May, the founder of U.S. server manufacturer Supermicro was indicted on charges of smuggling advanced chips into China.
“I’m a proponent of knowing customer laws regarding data centers around the world,” Bresnick said. “If you’re going to have a company conduct large-scale training on state-of-the-art hardware, you need a mechanism to report who they are and what they’re doing.”
President Joe Biden’s Commerce Department proposed federal know-your-customer rules for data centers in 2024, but there appears to have been no further progress under President Donald Trump’s administration. But exporters who ship advanced chips overseas are supposed to ensure they are used only for approved purposes.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
