Open Weights Are a Decoy. The Moat Was Never the Model.

In 2026 the price of a capable AI model fell to zero, and the price of everything required to run one at scale rose faster than almost any industrial input on earth. Those are not two stories. They are one story, and nearly every argument about open AI is being conducted on the wrong side of it.

On August 10, Meta released Muse Glimmer under Apache 2.0. Thirty billion parameters, quantized from over fifty five gigabytes down to under twenty, agentic function calling, multimodal input, a hundred languages. It runs on a single consumer graphics card or a recent MacBook, with the network switched off. Alongside it, Mark Zuckerberg published a six thousand five hundred word essay arguing that concentrating advanced AI in a few hands is the real danger, and that broad distribution is the safer path.

The industry spent the following day arguing about whether open weights are really open source, on the grounds that the training data and training code were withheld. That objection is correct and almost entirely beside the point. It is a dispute about the label on a box, conducted inside a building that is quietly changing hands.

Here is what is worth noticing instead. The layer that was given away is the only layer in the stack whose price was already collapsing. Every layer beneath it became scarcer, more expensive and more concentrated over the same period, and not by a little. By an order of magnitude, along a dimension that cannot be forked, downloaded, or licensed around.

This is not really about Meta. Meta has simply produced the cleanest illustration yet of something that has been true for roughly eighteen months.

What actually got cheaper

Glimmer is a distillation. It was compressed out of Muse Spark, a larger model that remains closed. That detail is usually filed under technical trivia. It is closer to a disclosure.

A distillation is what you produce when the expensive part is already finished and you want a cheaper copy of its behavior. The intelligence in Glimmer was paid for once, at frontier scale, on hardware nobody is giving away. What was released is the compressed residue of that spend. This is not a criticism of the technique, which is genuinely impressive engineering. It is an observation about which asset was actually transferred.

And the market price of that asset was already zero.

By the middle of 2026, Chinese open weight models accounted for roughly sixty one percent of all tokens consumed on OpenRouter. DeepSeek is the single largest provider on the platform by volume, ahead of Google, Anthropic and OpenAI. Qwen has crossed a billion cumulative downloads on Hugging Face, takes more than half of all open model downloads, and by February 2026 accounted for sixty nine percent of new derivative models. Llama has fallen off the OpenRouter rankings entirely, and its share of derivatives has gone from a quarter in late 2023 to about eleven percent.

Against that backdrop, Glimmer trades blows with a Qwen model that is one release from being superseded. This is not a knock on the engineering. It is a description of a commodity market. A thirty billion parameter model in August 2026 is undifferentiated, substitutable and price taking, which is the textbook definition of a commodity, and its price behaves accordingly.

Nobody gives away a moat. Companies give away inventory that has stopped appreciating. That is not cynicism, it is ordinary capital allocation, and it is the correct decision for anyone holding an asset the market has repriced to nothing.

What got scarcer

Now look one layer down.

Google, Amazon, Microsoft and Meta will spend roughly seven hundred and twenty five billion dollars in capital expenditure in 2026. That is up seventy seven percent from four hundred and ten billion the year before, itself a record. Add Oracle and the big five will put more than six hundred billion into infrastructure alone, around three quarters of it aimed at AI. To fund it, hyperscalers raised over a hundred billion in debt during 2025, with forecasts of well over a trillion in issuance ahead.

Those are staggering numbers, and they are still the least interesting part.

Because money is no longer the binding constraint either. Electricity is.

The Uptime Institute now identifies power as the single defining limit on data center growth worldwide. Gartner projects that by 2027, forty percent of existing AI data centers will be operationally constrained by power availability. The gating items are not chips. They are high voltage transformers, switchgear and grid interconnection. Lead time on a high power transformer was twenty four to thirty months before 2020. It now runs to five years, and tier one manufacturers are quoting sixty months and up. Interconnection queues in Northern Virginia, Phoenix and Dallas now run four to seven years, and something like two thousand gigawatts of generation and storage is sitting in the American queue waiting for access.

Sit with the shape of that. Not the size, the shape.

A model can be copied at the cost of bandwidth. A transformer cannot. There is no repository you can clone that shortens a five year manufacturing queue, no permissive license that moves you up an interconnection list, no community fork of a substation. The scarcest input in artificial intelligence is now a physical object with a half decade lead time, held in a queue, allocated to whoever placed an order and posted capital years ago.

The layer that was opened is infinitely reproducible. The layers that were kept are rate limited by heavy industry and regional utility planning. Those two facts are the entire argument.

Why the old intuition stopped working

The open source instinct is not naive. It was earned, and it was right, in a domain where the conditions genuinely held.

Linux mattered because once you had the source, you could run it. The artifact and the capability were the same object. A commodity server, a compiler, a copy of the code, and you were operating at the same tier as anyone else. Apache mattered for the same reason. The freedom was real because the marginal cost of exercising it was close to nothing. That is why open source reorganized the software industry rather than merely embarrassing it.

Models break that equivalence, and the break is not subtle.

Weights are the recipe. Serving is the kitchen. Downloading Glimmer gives you something real, and I want to be precise about that, because the on device story is genuinely significant and I am not dismissing it. You get a capable agentic model that runs offline, on hardware you own, under a license that lets you do as you please. On an RTX 5090 with speculative decoding, that is seventy five to two hundred and thirty tokens per second. That is a real capability, held by a real person, outside anyone's terms of service. It matters.

Now serve it to a million people.

The distance between "I can run this" and "I can serve this" is the entire industry. Open weights hand you the first with no conditions attached. The second is gated by capital you do not have, hardware allocations you cannot get, and grid capacity that takes five years to procure. We inherited an intuition from an era when possessing the artifact meant possessing the capability, and we never rechecked the premise when the premise changed.

So the freedom is genuine and the scope of it is small, and the gap between those two is being filled almost entirely with rhetoric.

Nobody planned this

It would be easier if this were a conspiracy. It is not, and pretending otherwise would get the mechanism wrong.

No one sat in a room and decided to stage a giveaway as misdirection. What happened is simpler and more durable than that. Every firm in this market is running the same unremarkable calculation. Open at the layer where you can no longer charge. Invest at the layer where scarcity is increasing. Do both at once, and describe the first one in the most flattering language available.

That is not villainy. It is what commoditization looks like from the inside, and any competently run company would do the same. Zuckerberg's essay may well be sincere. It can be entirely sincere and still be describing a layer that no longer determines the outcome. Both things fit comfortably in the same document.

The error is not theirs. It is ours. We keep reading a capital allocation decision as a moral one, and then feeling reassured by it.

What the word open now measures

There was a time when calling a technology open told you something about the distribution of power. That is what gave the word its charge. It implied that the thing which decided outcomes had been placed within reach of anyone who wanted it.

The word still carries that charge. It no longer carries that meaning.

Today, when someone tells you a model is open, they have told you something true and verifiable about a license and a file. They have told you nothing whatsoever about who controls artificial intelligence. Those two claims separated somewhere around the point where training runs started costing more than semiconductor fabs, and the vocabulary never caught up. So the word keeps doing emotional work it stopped being able to justify, and we keep accepting it as evidence about a question it cannot address.

This is why the open source purity argument is such an effective distraction, and I do not think anyone designed it to be. Whether Glimmer qualifies as open source under the strict definition, whether the training data should have shipped, whether Apache 2.0 is sufficient: those are all real questions, and every one of them is a question about the top of the stack. You can win that debate completely, extract full training data and full code from every lab on earth, and the concentration of power in AI will not have moved by a measurable amount. The constraint is not there. It is four layers down, in a queue for transformers.

The image to keep

Meta gave away a model that runs on a laptop with the wifi off. That is a real thing, and I would rather live in the world that has it than the one that does not.

Meta is also spending up to a hundred and thirty five billion dollars this year on the layer that is not for sale, in a market where the scarce input has a five year lead time and forty percent of the buildings will not be able to draw the power they were designed for.

We were handed the top floor. The foundation, the land, the substation and the five year queue for the equipment that energizes it belong to a handful of landlords who are, at this moment, buying more of it than anyone has ever bought of anything.

The keys are real. The building is not ours.


Shubhendu Tripathi is an AI and ERP strategy consultant based in Toronto, and the host of The Integration Layer, a podcast on AI, enterprise systems, and the work of making them fit together. Connect on LinkedIn or reach out at tripathis@qubittron.com.