← back_to_blog
// market data 6 min read

Open weights vs. frontier models: who actually holds the market in 2026

Two credible datasets put open-weight AI models at 11% and at roughly 60% of the market. Both are right. Here is what each one measures, why the gap matters for your network, and what it means for branch AI traffic.

Two horizontal bars comparing open-weight model share: 11 percent of enterprise LLM API usage versus about 60 percent of public router token volume.

Ask “how much of the AI market runs on open-weight models?” and you get two answers that are off by a factor of five. Both come from real data. They just count different things — and the difference is the whole story for anyone running a network that carries AI traffic.

  • 11% — open-source models’ share of enterprise LLM API usage, per Menlo Ventures’ 2025 State of Generative AI in the Enterprise.
  • ~60% — open-weight models’ share of token volume on OpenRouter as of mid-2026, up from roughly a third in late 2025.

Neither number is wrong. One measures what companies buy. The other measures what developers run. If you only track the first, you will be surprised by the traffic already crossing your branch links.

What enterprises buy: closed models, and more so than last year

Menlo Ventures’ 2025 report is the cleanest read on enterprise procurement. Its finding is blunt: open-source models hold 11% of enterprise LLM API usage, down from 19% a year earlier. Enterprise spend consolidated rather than diversified — Anthropic at 40% of enterprise LLM API spend, OpenAI at 27%, Google at 21%. Those three account for 88% of enterprise API usage, and foundation-model APIs pulled in $12.5B of the $37B enterprises spent on generative AI in 2025.

The reason is not benchmark scores. It is that a managed API comes with a support contract, an indemnity clause, an SLA, and nobody on your payroll paged at 2am because an inference server fell over. For a regulated buyer, that bundle is often worth more than the per-token savings.

What developers run: open weights, and increasingly so

Now look at the same question from the routing layer. OpenRouter’s 100-trillion-token study — covering November 2024 through November 2025 — found open-weight models at roughly a third of usage by late 2025, with proprietary models averaging about 70% of weekly token volume across the year. Chinese open-source models averaged around 13% of tokens, spiking to nearly 30% in individual weeks.

Through the first half of 2026 that trend accelerated sharply. Reporting on OpenRouter’s mid-June 2026 data puts open-weight models at roughly 60% of token volume — a reversal of the 60/40 split favouring proprietary models only three months earlier, driven largely by cheap, capable Chinese open models.

Different population, different answer. The router sees startups, indie developers, agent frameworks, and cost-sensitive workloads. The Menlo survey sees procurement departments. Your branch office contains both.

The capability gap is now months, not generations

The old argument for closed-only policy was that open models were simply worse. That gap has compressed to the point where it no longer settles the question on its own. Stanford’s 2026 AI Index put the leading closed model just 3.3% ahead of the leading open model as of March 2026, and Epoch AI measured an average four-month lag between open and closed releases over January–May 2026.

Four months is shorter than most enterprise procurement cycles. A policy written on the assumption that open models are a year behind is a policy written against last year’s facts.

Sovereignty is pushing inference back on-premises

Deloitte’s State of AI in the Enterprise 2026 surveyed 3,235 business and IT leaders across 24 countries and six industries. It found:

  • 83% view sovereign AI as important to their strategic planning
  • 77% now factor country of origin into vendor selection
  • Nearly 3 in 5 build their AI stacks primarily with local vendors

That is the demand side of the same force. When the model has to stay in a jurisdiction — or in a building — the deployment decision stops being about benchmarks and starts being about where the weights physically sit.

The money agrees. In four weeks during mid-2026, Fireworks, Baseten, and Together AI collectively raised about $3.8B to serve open-weight models as managed infrastructure, as Forbes reported. One detail from that piece is worth more than the valuations: Fireworks says more than 95% of the tokens it serves come from models specialised on customer data. That is not people picking open models to save money. That is people picking open models because the model has to be theirs.

What this actually means for your network

Put the three datasets together and the picture is not “open wins” or “closed wins.” It is both, permanently, on the same network:

  1. Your finance team’s copilot calls a frontier API in someone else’s cloud.
  2. Your engineering team runs an open-weight model on a box in the rack.
  3. Your field team’s agent calls whichever is cheaper this quarter.
  4. Somebody in marketing is pasting customer data into a model you have never heard of.

Three of those four are invisible to a network built for the pre-AI era. They are all just HTTPS on 443. The mix between them shifted by 30 points in six months on the public routers — which means any control you build around a fixed list of approved providers will be out of date before the fiscal year ends.

The control that survives the mix changing is control at the wire: policy applied to AI calls as they leave the branch, regardless of which model is on the other end. That is what the AI gateway built into our branch appliance does — virtual API keys per team, per-key model allow-lists and budgets, an audit trail of every call, and PII redaction before the request reaches a provider. Route a call to a local model when the data cannot leave, and to a frontier API when it can, under one policy.

It is the same idea as the rest of the fabric: bond the links, mesh the branches, govern the tokens, from one appliance and one dashboard. AI governance is one of six jobs that appliance does — it did not need a new box, because the traffic was already passing through this one.

A short checklist

If you do nothing else this quarter:

  • Measure the split. You cannot govern a mix you have not counted. Find out what share of your AI calls leave the building today.
  • Assume both. Write policy that names data classes, not vendors. “PII never leaves the branch” survives a model swap; “only OpenAI is approved” does not.
  • Put the enforcement point where the traffic is. At the branch edge, not in a portal that a determined user can route around.
  • Give teams a sanctioned local option. Shadow AI is what happens when the approved path is slower than the unapproved one.

Caged AI is built by Mushroom Networks, which has spent two decades shipping SD-WAN and broadband bonding into broadcast trucks, hospital networks, and government agencies. The AI part is new. The part where we care what crosses the edge is not.

Want the numbers for your own sites? Get more info, or if you resell and manage networks for other people, look at the partner program.


Sources

Figures are as reported by the sources above at the dates given. The mid-2026 OpenRouter token-share figure comes from press reporting on OpenRouter’s June 2026 data rather than from a primary release, and is described here as approximate.

// 04 — pilot

Run a pilot at
one branch.

Drop a Caged AI box at a single site. Bring your existing WAN links. Bring your existing AI providers. Replace it inside a month — or keep going.