On September 8, Treasury Secretary Scott Bessent told an audience at SMU’s business school: “The technical word for stealing and copying American AI models is distillation. So, the Chinese distil our models and they can never get ahead of us.” His analogy: “if you’re looking over someone’s shoulder, copying their homework, you can never get a higher grade than they do.”

The same day the NSA, CISA and FBI published AA26-251A, a joint advisory naming six Chinese labs. Four days after that, Dario Amodei’s pacing essay linked the advisory in its China section and asked for a crackdown. I wrote that post about the open weights the essay never names. This one is about the sentence it depends on.

What the Advisory Says

The advisory is behavioural, not forensic. No IPs, no hashes. It reads like a threat-intel report written by people who have seen the labs’ internal data and cannot publish it.

  • Who. “DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024.” Claude and GPT appear on all six target lists.
  • How. “A gray market of API proxies known as ‘transfer stations’ to bypass U.S. AI companies’ regional restrictions, breach terms of use, evade safeguards, and undermine traceability.” Fraudulent accounts, shared payment methods, 24/7 usage with no idle periods, new accounts hitting maximum throughput on day one.
  • The DeepSeek line. “DeepSeek’s publicly quoted training costs of $5.6M are misleading as it does not include the true cost of the data acquired through extensive malicious distillation.” That is the US government taking a position on the number that moved Nvidia’s share price in January 2025.
  • The recommended countermeasure is the interesting one: “Subtly alter responses for suspected malicious distillation attempts to attenuate the payoffs.” Poison the well for the accounts you suspect. That is a product decision, made by a security agency, about what a model should say.

The Beam

Put Bessent’s sentence next to the pacing plan and the dependency is visible.

Amodei’s essay proposes slowing US frontier capability by regulation and voluntary agreement, and treats China separately: chips, distillation, weight theft. He says Anthropic has “always understood that they would be essential to any pacing.” He is right, and the reason is arithmetic. A pace that binds US labs costs the US its lead unless China’s rate of progress is bounded by the US rate. Distillation is the mechanism that makes it bounded. If the only way to catch up is to copy, then slowing the copied slows the copier, and pacing is free.

That is why the Hacker News reply to the essay that mattered was three words long: “China will also slow down”, with a link to the advisory. And it is why the government’s advisory landed four days before the essay rather than four days after. The crackdown is not a side policy. It is the load-bearing assumption.

The Chinese distil our models and they can never get ahead of us.

— Scott Bessent, SMU Cox School of Business, September 8 2026

The Load

Anthropic published its September threat report two days after the advisory. It is the first time the scale has been put in one table, and the direction of the numbers is the problem for the beam.

  • February. Anthropic’s first distillation post: three labs, “over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts.”
  • May to July, Alibaba alone. “Peaked at nearly 3 million exchanges per day launched from more than 3,500 fraudulent accounts.” Total: “over 151 million exchanges observed.” The target was chain-of-thought traces for agentic tasks, software engineering and kernel development.
  • The ban cycle. “The first consisted of nearly 5,000 fraudulent accounts leveraging residential proxies, disposable emails, and virtual-card payments… When we banned this pool of accounts, Alibaba quickly shifted its traffic through the second pool.” Some of those accounts were also carrying DeepSeek and Xiaomi traffic. Shared infrastructure, shared resilience.
  • The relay twist. Moonshot “silently forwarded customer requests to Claude, instead of processing them using Kimi”, almost 300,000 requests in ten days through 5,380 accounts, then built “a CoT extraction pipeline” on the saved transcripts. DeepSeek did the same. Xiaomi replayed its own users’ coding sessions through Claude, 400,000 requests across 1,500 accounts. Three labs sold Claude under their own brand and kept the receipts.

Nine months, one order of magnitude, and a ban that buys days. The advisory’s own detection guidance is usage-pattern analysis, which is what Anthropic was already doing when Alibaba moved pools. The beam is holding weight. It is not holding it well.

The homework analogy has a hole, and it shipped this week

Copying caps your grade only if you stop at copying. On September 10 Cognition shipped SWE-2, a reinforcement-learning post-train of Moonshot’s Kimi K3. On Cognition’s own table it scores 92.8% on Terminal-Bench 2.1 against Fable 5.1’s 91.4%. It loses badly on Terminal-Bench 4, and the numbers are the vendor’s, but the mechanism is the point. A student who copies the homework and then trains on the answers can beat the original on some tests. Distillation is a floor, not a ceiling, once you add your own RL. The “never get ahead” claim is true of the copy and false of what gets built on it.

What I Got Wrong in July

When Kratsios accused Moonshot of distilling Fable to build Kimi K3, I leaned on a persona-probing study suggesting K3’s Claude-like behaviour looked more like training-data contamination than targeted extraction, and quoted Dean Ball saying performance “can’t be explained away by distillation”. The September report changes the evidence. Anthropic now describes a Moonshot chain-of-thought extraction pipeline and 23 million exchanges between May and July, which is K3’s training window. I still think K3’s capability is mostly Moonshot’s own work, and the report does not claim otherwise. But “was K3 distilled from Claude at all” now has a documented yes attached to it, and I was too quick to file it under theatre.

What This Doesn’t Settle

  • Beijing says none of it happened. The Commerce Ministry called the advisory “groundless and without legal basis” and said the US “politicizes distillation, which is a normal technical and commercial issue”, with “resolute countermeasures” if sanctions follow. None of the six labs commented. No Entity List addition has landed. Bessent’s “on the table” is where it was in July.
  • The advisory shows no evidence. It asserts and maps to MITRE ATLAS. The numbers come from Anthropic, one of the companies the advisory is written on behalf of. I believe them because the account-level detail is specific and falsifiable. Vet the incentive anyway.
  • Distillation from outputs is not scraping, and it is not theft either. I made that distinction in February and stand by it. The Hacker News comment on the advisory writes itself: “US-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against Content Creators And Copyright Holders World Wide.” The terms-of-service case is strong. The moral high ground is contested.
  • The timing is diplomatic. Xi proposed a BRICS “open-source zone for artificial intelligence” on September 13. Trump hosts Xi on September 24. An advisory, a Treasury speech and a CEO essay in one week before a summit is a negotiating position, whatever else it is.

The Takeaway

  • Pacing is free only if China can only copy. That is the assumption under Amodei’s plan, and Bessent said it out loud four days earlier.
  • The copying is growing, not shrinking. 16 million exchanges in February. 151 million from one lab by July. Bans buy days.
  • The copy is a floor. SWE-2 on Kimi K3 beat Fable 5.1 on one benchmark this week. Homework plus training beats homework.
  • Watch September 24. If distillation is on the summit table, the beam gets tested in public. If it is not, the advisory was for the essay.