The Safest Small Model Is the One You Distilled Yourself. Everyone Is Doing the Opposite.

Enterprises are moving real workloads onto small models, and most are picking the option with no chain of custody. Distilling your own small model from a frontier model you already pay for is expressly addressed, in writing, by the companies that own those frontier models. Downloading an open-weight model whose lineage is the subject of letters to US senators is neither permitted nor forbidden. It is undocumented, and undocumented is the condition you will actually be asked about.

I am not arguing against small models. The economics are real and the privacy case is genuine. I am arguing that the two paths to a small model have completely different paperwork, that almost nobody compares them on that basis, and that the path most teams chose because it felt safer is the one they cannot evidence.

You moved the weights. You did not move the provenance.

The reasoning that gets a private model approved is familiar. Frontier APIs mean your data leaves the building. Self-hosting an open-weight model means it does not. Gartner expects organisations to use small, task-specific models three times more than general-purpose ones by 2027, driven mostly by cost and privacy.

Every part of that reasoning is about where the data goes. None of it is about where the model came from.

That was a fine trade when a model was a piece of software. The interesting question about a small model is not what it does, it is what it was taught from and by whom. You took custody of the weights. The provenance stayed upstream, with a party you have no contract with.

What the paperwork actually says

Here is the part that surprises people: distilling your own model is not the forbidden option. It is the documented one.

Anthropic's policy page on the question, dated 16 March 2026, is direct, and its baseline is the stricter of the two majors. It prohibits using its services to train or develop AI models without written permission, then grants that permission back for the non-competing case: you can use Claude's Outputs to train models that do not compete with Anthropic's own. It names the permitted categories, and they are ordinary enterprise work: sentiment analysis, content categorization, summarization, information extraction, semantic search, anomaly detection. What stays prohibited is the competitive case, general purpose chatbots and models built for open-ended generation, along with reverse engineering the training methods. The stated reason is that safety controls can be lost when a model trains on Claude's outputs without oversight.

OpenAI draws its line in contract language instead. In the current Services Agreement, section 3.3(e), the customer will not, "except for a Permitted Exception, use Output to develop artificial intelligence models that compete with OpenAI's products and services." A Permitted Exception is then defined as using Output to "develop artificial intelligence models primarily intended to categorize, classify, or organize data (e.g., embeddings or classifiers), if these models are not distributed or made commercially available to third parties," and to fine tune or customize models offered as part of OpenAI's own services.

Read those two side by side, because they are not the same permission.

The asymmetry nobody checks

Anthropic names summarization and information extraction, with no condition about who the resulting model is offered to. OpenAI's Permitted Exception is narrower on both axes at once. It covers models "primarily intended to categorize, classify, or organize data," which is a smaller set than summarization, and it alone attaches the condition that the model must not be distributed or made commercially available to third parties.

So consider an ordinary enterprise project. You distil a small model that reads incoming supplier documents, extracts the fields, and writes a short summary for an approver. Purely internal, it sits comfortably inside Anthropic's named list and inside OpenAI's exception. Now your product team wants to offer the same capability to your customers as a feature. Under Anthropic's wording you are arguing about whether it competes. Under OpenAI's agreement you have just stepped outside the express safe harbour, because not being offered to third parties was part of what the exception covered.

Be precise, because this is easy to overstate. OpenAI's restriction only ever bites on models that compete with its products and services, and the Permitted Exception is a safe harbour inside that restriction rather than the outer edge of what is allowed. Leaving it does not put you in breach. It removes your certainty, at exactly the moment the project stops being internal.

Same model, same training data, same use case, materially different amounts of contractual certainty, and the only variable is which API you pointed at it eighteen months ago.

This is the practical failure. Distillation does not feel like a licensing decision, it feels like engineering. So it gets made by a team that has never read the agreement, against a vendor chosen for latency and price, and nobody writes down which side of the line the project sits on while the answer is still cheap to establish.

What you actually downloaded

Now the other path. When you pull an open-weight small model, you are not getting a model that learned its capability from scratch. Distillation is the standard way the small members of these families are built, and the labs say so in their own reports.

DeepSeek distilled R1's reasoning traces into Qwen and Llama students. Qwen3 used strong-to-weak distillation from large teachers at roughly a tenth of the GPU hours its reinforcement learning pipeline would have needed, with better results. Gemma 3's post-training relies on an improved version of knowledge distillation from a large instruction-tuned teacher, and Gemma 4 follows the same recipe. DeepSeek V4 describes a multi-teacher pipeline in more detail than anyone else has.

None of that is scandalous. It is the state of the art, and the licences on those weights are permissive. The point is narrower and it is about evidence: the capability in the model you downloaded was transferred from a teacher, and you cannot name the teacher. That is not a criticism of the model's quality. It is a description of your documentation.

Two paths reach the same small model. Only one of them can be evidenced end to end.

The disclosures that made this a live question

Through 2026 the teacher question stopped being academic.

On 23 February 2026 Anthropic published a report alleging industrial-scale distillation campaigns by three named labs: DeepSeek, Moonshot AI and MiniMax. It said it had identified roughly 24,000 fraudulent accounts generating more than 16 million Claude exchanges, aimed at extracting agentic reasoning, tool use and coding: MiniMax above 13 million of them, Moonshot above 3.4 million, DeepSeek above 150,000. Days earlier, OpenAI had told the House Select Committee on China that DeepSeek employees used third-party routers to get around its access restrictions and wrote code to harvest outputs for distillation programmatically.

It kept going. On 23 April the White House Office of Science and Technology Policy issued a memo, signed by its director Michael Kratsios, accusing entities principally based in China of deliberate, industrial-scale campaigns to copy American frontier systems. On 30 April, testifying in his own lawsuit against OpenAI, Elon Musk conceded that xAI had in part trained Grok on OpenAI's models, describing it as general industry practice. Then in a letter dated 10 June to the Senate Banking Committee, Anthropic alleged that operators linked to Alibaba's Qwen lab had run nearly 25,000 fraudulent accounts and 28.8 million Claude exchanges between 22 April and 5 June, more than the three February labs combined.

Now the part that matters more than any individual allegation.

None of it is settled. No litigation has been filed over any of the distillation allegations against these labs. Moonshot has publicly denied a later White House accusation and credited its own architecture, and independent researchers were openly sceptical of that particular claim on timing grounds. Alibaba has not answered it, though it banned Claude for its own employees in July. DeepSeek and MiniMax have said nothing. Washington, meanwhile, has begun moving from letters toward instruments, with sanctions and Entity List designations raised by the Treasury Secretary in July and the bipartisan BLADE Act introduced on 5 August.

You do not have to believe any of the allegations. That is the point, and it holds either way. DeepSeek and Qwen are among the most widely deployed open-weight families in enterprise infrastructure, chosen overwhelmingly by teams whose stated reason was reducing risk. A dispute you cannot adjudicate is still a dispute you cannot document.

Why this lands on you rather than on the lab

You might reasonably say this is the labs' argument, not yours. You might add that the deadline moved, which it did.

Article 10 of the EU AI Act sets data governance requirements for high-risk systems: the origin of the training data and how it was collected, the preparation steps applied to it, and the quality criteria it has to meet. Those obligations were due on 2 August 2026. They are not in force. The Digital Omnibus, Regulation (EU) 2026/1744, entered into force on 27 July and moved standalone high-risk systems to 2 December 2027, with embedded systems following in August 2028.

I wrote about that deferral while it was still a proposal and my read has not changed. A delay on a records requirement is not relief, because the records describe work you are doing now. The trace you cannot produce in December 2027 is the one you did not capture in 2026.

Note also who owes it. Article 10 binds providers, not deployers. Buy a vendor's system and use it and your obligations sit in Article 26, not here. But fine-tune an open-weight model and put the result into service under your own name for a high-risk purpose and you are the provider of that system. You built it, so the duty is yours.

Then the gap. A provider's technical documentation, the Annex IV set that Article 11 requires, normally leans on what the upstream model provider passes down. That is precisely the duty the open-source exemption removes: Article 53(2) frees free and open-source model providers from those technical documentation obligations unless the model carries systemic risk. The copyright policy and the public summary of training content survive; the document you would have leaned on does not. The relief is theirs. The obligation stays yours.

And the regulator is not first in the queue anyway. Your own enterprise customers ask sooner, through a supply-chain questionnaire, which has no penalty tier, no grace period and no December 2027. It just has a field you leave blank.

The Chain of Custody Test

Five questions before a model goes into a process that matters. They take an afternoon while the model is still a candidate, and they are close to unanswerable a year after go-live.

Can you name the teacher? Not the model family, the specific system whose outputs or logits produced the capability you are relying on. If the honest answer is that it is somebody else's teacher and you do not know which, that is your finding.

Do you hold the permission in your own name? A permissive weights licence covers the weights. It does not stand in for a licence to the training signal upstream of them. When you distil from an API you pay for, the permission sits in your agreement, addressed to you.

Can you produce the traces? Retained, versioned, and tied to a specific model version. This is the artefact the Annex IV technical documentation will ask for, and the only path that reliably generates it is the one where you did the distilling.

Is your use written down on the permitted side of the line? Not assumed. Written, against the specific clause, including whether the model will ever be distributed or made commercially available to third parties, because for at least one major vendor that single condition decides whether you sit inside an express safe harbour or go back to arguing about competition.

Who answers if the lineage is later held defective? For a model you distilled under contract, there is a counterparty. For weights you downloaded, the answer is you.

The honest counterargument

The strongest objection is that this is risk theatre. Open weights carry permissive licences, no distillation case has been filed, cross-border enforcement is genuinely hard, and the cost gap is real. A team that self-hosts a capable open model is getting a working system today at a fraction of the API cost.

That objection is largely correct, and it answers a question I am not asking. This is not a prediction that you will be sued. It is an observation that you will be asked to attest, and attestation arrives years before litigation does. The questionnaire, the auditor and eventually Annex IV all want the same artefact, and none of them are waiting for a court.

The second objection is better, and it is technical. Distilling your own model is not automatically the efficient choice. A compute-equivalent study by Boizard and colleagues found that at matched training and inference FLOPs, ordinary instruction fine-tuning sits on or near the Pareto frontier across most configurations, and that reasoning distillation reaches the frontier only on open-ended tasks at 7B parameters and above. A sequential curriculum using just 25 to 50 percent reasoning data captured most of the benefit at a fraction of the cost. So the honest version of my recommendation is not that distillation is free or always right. It is that when you do build a small model, building it from a teacher you have a contract with costs little more than building it from one you do not, and it is the only version that leaves a record.

Where this lands

Do not stop using open weights. That is not the conclusion, and the cost and privacy arguments that got you here still hold.

The conclusion is that the two paths to a small model were never equivalent on the axis you will be tested on, and the enterprise picked the undocumented one while believing it was the cautious one. Distilling into a narrow, task-specific model, from a frontier system you already license, is addressed in writing by the vendors, produces the provenance record the regulation will ask for, and gives you a counterparty when someone asks a hard question.

Build the derivative. Do not inherit one.


Shubhendu Tripathi is an AI and ERP strategy consultant based in Toronto, and the host of The Integration Layer, a podcast on AI, enterprise systems, and the work of making them fit together. Connect on LinkedIn or reach out at tripathis@qubittron.com.