AI Progress from BERT to Frontier Agents: Capability-Cost Collapse and Specialized Models
A recent paper on arXiv (2608.13675) details the swift advancements in AI from October 2018 to July 2026, emphasizing a significant decline in the capability-cost curve. It highlights the evolution from basic models like BERT to advanced agents adept at tackling intricate mathematics and programming. Since late 2024, the capacity to address genuine coding challenges has surged nearly sixfold annually. Costs have drastically decreased: OpenAI's GPT-5.6 Luna achieves leading capabilities for merely $1–6 per million tokens, significantly cheaper than previous iterations. Specialized models now dominate performance: Claude Opus 5 excels in frontend development, Claude Fable 5 is best for repository coding, and GPT-5.6 Sol leads in terminal operations. In a grade school math assessment with the Qwen 2.5 model, basic methods answered 58 of 100 problems, while advanced techniques solved up to 79. A confidence ranking tool successfully identified 47 correct answers among its top 50 selections, proving effective for task sorting. All research materials have been made publicly accessible.
Key facts
- Paper arXiv:2608.13675 covers AI progress from October 2018 to July 2026.
- AI models evolved from BERT to agents solving complex math and writing software.
- Real coding issue resolution improved nearly six times per year since late 2024.
- OpenAI's GPT-5.6 Luna matches flagship capabilities at $1–6 per million tokens.
- Claude Opus 5 leads in frontend coding.
- Claude Fable 5 excels at repository-level coding.
- GPT-5.6 Sol dominates terminal tasks.
- In a grade school math test, Qwen 2.5 basic methods solved 58/100, advanced sampling up to 79.
- Confidence ranking tool identified 47 correct answers in top 50 choices.
Entities
Institutions
- OpenAI
- Anthropic
- Alibaba
- arXiv