Benchmark Trends: Open Source LLM Gap to AI in Math Research
연구/벤치마크 | Sat Jun 27 2026 00:00:00 GMT+0000 (Coordinated Universal Time) | 3 sources
Doubleword analyzed the gap between open weights and closed source LLMs, Terence Tao envisioned AI-driven Big Mathematics, and the Anthropic Economic Index measured usage patterns.
Analysis
[Doubleword] published analysis of the gap between open weights and closed source LLMs [1]
- Analysis based on the Artificial Analysis Intelligence Index
- Single metric predicts zero-month gap by December 2026
- Average across 18 benchmarks maintains roughly a 5-month gap
- Gap on coding benchmarks shrank sharply from 15 months to 1-2 months
[Terence Tao] presented a vision for the AI-driven era of Big Mathematics [2]
- Concept proposed by UCLA mathematician Terence Tao
- Solving complex mathematical problems through human-machine collaboration
- Raises possibility of fundamental changes in mathematical research methods
[Anthropic] released the Economic Index June 2026 report [3]
- Reflects increased use of long-running agentic tasks via Claude Code and Cowork
- Introduced high-frequency sampling trackable down to the hour
- Applied a new classifier to label conversation outputs
- Separately released Claude conversations and 1P API results
[Anthropic Economic Index] measured daily rhythms of AI usage and economic impact [3]
- Work-related queries decline on weekends
- with less decline among high-income professions
- News requests peak in the morning
- sleep advice requests peak around 5 AM
- Tax-related requests surge around tax filing deadlines
- Greater compute usage correlates with more valuable outputs