← Back to all insights

Technology Published · 15 September 2026

What frontier models can actually do in 2026

Capability has climbed steeply in twelve months, but unevenly. Stanford's AI Index carries an uncomfortable finding: the same model that solves competition mathematics fails to read an analogue clock.

6 min read

If you made an AI decision a year ago based on what models could do then, that decision is stale. If you made it based on what they couldn’t do, it is probably staler still.

The twelve-month jump, in numbers

The 2026 AI Index from Stanford HAI — its ninth annual edition, and the sector’s reference measurement — tracks progress against stable benchmarks. Two figures sum up the year:

  • On SWE-bench Verified, which measures resolution of real GitHub repository issues, performance rose from 60% to near 100% in a single year.
  • On OSWorld, which evaluates agents operating an actual computer, task success climbed from 12% to roughly 66%.
SWE-bench Verified — real GitHub issues 60% ~100% OSWorld — agents operating a computer 12% 66% PREVIOUS YEAR AI INDEX 2026
AI Index 2026 (Stanford HAI), chapter 2. For contrast: the same models read an analogue clock correctly 50.1% of the time.

The second number is the one that should change your planning. A 12% success rate is a lab curiosity. 66% is something you can put behind a process, provided you design for the other 34%.

Capability is jagged, not uniform

Here is the part almost nobody mentions. The same report finds that top models read an analogue clock correctly just 50.1% of the time. A task any eight-year-old handles.

This is not a funny aside. It is the single most important property of the technology you are working with. Capability in these systems is not a line going up; it is a serrated profile. They are superhuman at dense symbolic reasoning and mediocre at things that strike us as trivial because our visual perception is hardwired for them.

The practical consequence: you cannot extrapolate. A model passing PhD-level science questions tells you nothing about whether it will parse your supplier’s scanned delivery note. The only way to know is to test it on your data. Every time.

The top of the market has flattened

As of March 2026, frontier models from Anthropic, xAI, Google, OpenAI, Alibaba and DeepSeek sit clustered within 25 Elo points of each other on the Arena leaderboards. The leading US model’s edge over Chinese competitors is 2.7%.

Translated for anyone buying technology: model choice has stopped being a strategic decision. When six vendors are within a hair of each other, competition shifts to cost, latency, reliability and data processing terms. Which are, as it happens, the criteria that should always have governed.

It also means locking yourself to one vendor makes less sense than ever. Design the integration so switching models is a configuration change, not a project.

Who is building this

More than 90% of notable frontier models in 2025 came out of industry, not academia. US private AI investment reached $285.9 billion in 2025.

Worth holding in mind when you assess a vendor’s staying power: that spending pace is unsustainable for most players. Some of today’s names will not exist in three years, and your architecture should outlive that.

What to do about it

  • Revalidate what you ruled out a year ago. Use cases you rejected because “the model isn’t there yet” deserve a second test. Many are there now.
  • Don’t extrapolate from benchmarks to your business. Build an evaluation set from your own cases before committing to anything.
  • Treat the model as a commodity. Competitive advantage sits in your data and your workflow, not in which API you call.
  • Design for failure. At 66% agent task success, the error path is the product.

Capability has stopped being the bottleneck in most projects we see. The bottleneck is knowing what to ask of it.

Sources

Next step

How ready is your business for AI?

Evaluate your AI maturity in 5 minutes and get free personalised recommendations.

Ready to move beyond the hype?