← Back to all insights

Technology Published · 28 August 2026

Open weights versus closed models: the capability gap closed, the adoption gap didn't

In March 2026 the best open model landed 49 Elo points behind the best closed one. A 3.4% gap. And yet open-source share in enterprise deployments fell from 19% to 11%. That paradox has an explanation.

6 min read

For three years the argument against open-weight models was simple: they are worse. That argument no longer holds, and yet enterprise adoption has moved the other way.

The numbers

The 2026 AI Index from Stanford measured the distance in March 2026: the leading closed model scored 1,503 on Arena and the leading open-weight model 1,454. Forty-nine Elo points — a 3.4% gap.

Leading closed model 1,503 Leading open-weight model 1,454
Arena score, March 2026. AI Index 2026 (Stanford HAI). Bars start at zero: the near-identical lengths are the finding.

A gap once measured in years is now measured in low single-digit months. DeepSeek, Llama, Qwen and Mistral closed it faster than almost anyone predicted in 2025.

And here is the odd part: open-model share in enterprise deployments fell from 19% to 11%.

Why adoption falls while capability rises

It is not irrational. Capability was never the real criterion.

What you buy in a closed model is not intelligence, it is indemnity. A contract, an SLA, someone to call when it breaks, a data processing commitment your legal team can file. Running open weights means you take on that role yourself.

Cost shifts, it does not vanish. The model is free; the GPUs, the team who knows how to operate them, the on-call rota and the upgrade cycle are not. At medium volumes the arithmetic almost always favours the API. The crossover point arrives later than people think.

The refresh pace is brutal. Adopting open weights is a commitment to re-evaluate every few months. Plenty of teams would rather that were somebody else’s problem.

When it does make sense

Three cases where we recommend open weights without hesitation:

  • The data cannot leave. Medical records, case files, trade secrets. If moving data off your network is a flat no, the conversation ends there.
  • Very high volume, very narrow task. Classifying, extracting, routing, tagging. Millions of calls a month for something an 8B model handles. Self-hosted inference is orders of magnitude cheaper there.
  • Latency or disconnection. A plant without guaranteed connectivity, a field device, millisecond requirements.

Outside those three, starting on an API and migrating later is usually right — and it is reversible, which is the point.

The part nobody mentions

The capability gap closed; the safety gap did not. Open models arrive without the guardrails closed labs apply at the service layer, which makes you responsible for building them: input and output filtering, rate limits, logging, abuse controls.

If you are deploying open weights in front of end users, budget for that layer. It is not optional and it is not free.

The sensible call

Don’t pick a side. Most architectures we see working well use both: a small self-hosted model for the 80% of traffic that is high-volume and low-reasoning, and a frontier API for the 20% that genuinely needs capability.

What you need to design is not the choice, it is the ability to change your mind. A thin abstraction over the provider, your own evaluation set, and the discipline to re-measure quarterly.

With a 3.4% difference and a market that moves monthly, locking yourself to either side for three years is the one clearly wrong decision.

Sources

Next step

How ready is your business for AI?

Evaluate your AI maturity in 5 minutes and get free personalised recommendations.

Ready to move beyond the hype?