Open AI Models Are Catching Up Twice as Fast Each New Era
A SemiAnalysis analysis finds open-weight AI models now close the gap with the best closed models roughly twice as fast with every new era of development — from 19.7 months down to 4.8 months.

Open-weight AI models are catching up to the best closed models roughly twice as fast with each new era of large language model development, according to a SemiAnalysis analysis titled "Are Open Models Catching Up?", published August 21, 2026. Rather than judging every model on one fixed benchmark suite that saturates over time, SemiAnalysis compares each era using the benchmarks that actually defined it.
Three eras, three catch-up times
The first era, early scaling, ran from 2022 to 2024. Its closed-model reference was OpenAI's GPT-3.5 Turbo, released in November 2022, evaluated on GSM8K, HumanEval, TriviaQA and MMLU-Pro. Meta's Llama-2-70B was the first open model to come close, scoring 39.9 on a composite scale against GPT-3.5 Turbo's 75.7. The gap closed only when Meta shipped Llama-3.1-405B in July 2024, scoring 86 — a full 19.7 months after GPT-3.5 Turbo's release.
Later in the same era, DeepSeek V3 matched OpenAI's GPT-4o in December 2024, scoring 94.1 against 95.5, and Alibaba's Qwen2.5-72B came close with a model six times smaller than Llama-3.1-405B — 72 billion parameters instead of 405 billion — after being pretrained on 18 trillion tokens.
The second era, reasoning, ran from 2024 to 2025. Its reference was OpenAI's o1, released September 12, 2024, evaluated on GPQA-Diamond, AIME 2026, SimpleQA Verified and Humanity's Last Exam. DeepSeek-R1-0528 closed an initial 12.1-point gap in May 2025, reaching a score of 78 against the target — in just 8.5 months.
The third era, agentic, runs from 2025 to today. Its reference is Anthropic's Claude Opus 4.5, released in November 2025, evaluated on complex agentic tasks such as coding and computer use, using Terminal-Bench 2.1, BrowseComp-Plus, τ³-Banking and DeepSWE. Moonshot AI's Kimi K2.6, released in April 2026, closed the gap with Claude Opus 4.5 in just 4.8 months — the fastest catch-up SemiAnalysis has observed to date.
The catch-up time keeps halving
- Early scaling era — GPT-3.5 Turbo to Llama-3.1-405B: 19.7 months
- Reasoning era — o1 to DeepSeek-R1-0528: 8.5 months
- Agentic era — Claude Opus 4.5 to Kimi K2.6: 4.8 months
Across the three eras, both the initial score gap between the first closed model and the best open model, and the time needed to close it, keep shrinking — the catch-up time has roughly halved with every new era. SemiAnalysis is careful to flag real limits to this comparison: benchmarks are an imperfect proxy for real-world work, and model developers can deliberately optimize a model for a specific benchmark rather than for general capability.
Why this threatens closed labs' margins
If open-weight models can match frontier closed models for a fraction of the cost, the model layer risks becoming commoditized. That is a serious threat to the margins of frontier labs such as OpenAI and Anthropic, which have raised enormous sums of money to build closed models on the bet that capability, not just price, would keep them ahead of open alternatives.
A wider race: the US-China backdrop
GIGAZINE, reporting on the SemiAnalysis analysis, adds that Bloomberg has separately tracked a similar dynamic in the broader US-China AI race, with China sometimes pulling ahead on usage and cost even where it trails on raw capability. GIGAZINE also notes that in July 2026, several US AI startups sent an open letter to the Trump administration asking it not to ban Chinese open-weight models, arguing that access to those models benefits American developers too. This is one link in the agentic era SemiAnalysis describes, since agent deployments are exactly where Terminal-Bench-style evaluations, one input SemiAnalysis relies on for its era-3 comparison, are most heavily used.
For enterprises evaluating AI vendors today, the trend matters beyond the benchmark scores themselves. If a credible open-weight alternative now appears within months, not years, of a new closed-model release, teams planning a 2026 or 2027 deployment have a wider field of vendors to test and negotiate with before committing budget to a single API, and less reason to assume that today's best closed model will remain unrivaled by the time a project ships.
Sources
- Are Open Models Catching Up?SemiAnalysis · August 21, 2026
- AI分野では「オープンモデルがクローズドモデルに追いつく時間」が飛躍的に縮まっているとの指摘GIGAZINE · August 24, 2026



