A new empirical review tracks how AI models advanced from BERT in October 2018 to frontier agents that write software and solve math by July 2026 arXiv DOI. The authors posted the study to arXiv in August 2026 arXiv DOI and released all research materials publicly. The work is an empirical review with reproducible figures, not a new model release.
The headline shift is in coding. The review finds the ability to resolve real-world coding issues improved roughly sixfold per year since late 2024 arXiv review, a pace the authors say outstrips earlier model generations.
Cost moved the opposite direction. The paper reports OpenAI’s budget model GPT-5.6 Luna matches flagship-level capability for about one to six dollars per million tokens arXiv review, undercutting older versions that cost far more for weaker output. The authors frame this as the collapse of the capability-cost curve, where price no longer tracks a model’s performance tier.
Performance leadership has also fragmented. The review places Claude Opus 5 in front-end coding, Claude Fable 5 in repository-level work, and GPT-5.6 Sol in terminal tasks arXiv DOI, so buyers now choose by task rather than by a single brand hierarchy.
On a grade-school math test using Qwen 2.5, basic methods solved 58 of 100 problems while advanced sampling reached 79, and a confidence-ranking tool flagged 47 correct answers in its top 50 picks arXiv review.
For builders, the practical move is portfolio thinking: match the model to the job and to a budget band, because the frontier is no longer one model. The full dataset and figures behind the review are public. Related reading: arXiv’s agent-memory model now bars stale, retracted data zBrandco and a separate arXiv study tracks edge agents cutting latency violations zBrandco.
