Expanding AI Research Ecosystem and New Trends in Benchmark Experiments
연구/벤치마크 | Thu Jul 23 2026 00:00:00 GMT+0000 (Coordinated Universal Time) | 7 sources
AI research and evaluation is diversifying through OpenAI's national science support, Anthropic's economic research fund, and pelican benchmark experiments.
Analysis
[OpenAI] announced US national science infrastructure support program [1]
- Provided $4M in Codex access to approximately 2
- 000 researchers participating in the Genesis Mission
- $3M API support for large-scale scientific campaigns
- Up to $10M API usage available with $2.5M spend
- Opened GPT-Rosalind bioscience-specialized model to national laboratory researchers
[Anthropic] unveiled research agenda for $200M Economic Futures Research Fund [2]
- External research support to prepare for AI's economic impact
- Selected 5 priority research areas (AI impact in the workplace
- transition support
- income support modernization
- etc.)
- Support centered on large-scale RCTs and creative pilot programs
- Shift from existing small grants to large-scale projects
[Anthropic Economic Index] launched Economic Index connector for Claude [3]
- Enables direct exploration of AI usage data within Claude
- Supports queries on AI usage patterns by occupation
- region
- and task
- Activated in claude.ai connector directory
- no installation required
- Compatible with all Claude models
[Pelicanmaxxing Experiment] released experiment verifying whether AI labs overfit to benchmarks [5]
- Tested 1
- 008 SVG generations across 7 frontier models
- Designed 48 prompts combining 8 animals × 6 vehicles
- Used GPT-5.6 Luna as judge and Gemini 3.1 Flash-Lite for feature extraction
- Verified whether the pelican-bicycle cell significantly outperforms other combinations
[Cactus Hybrid] released Gemma 4-based on-device confidence routing model [6]
- Embedded a probe inside the checkpoint that returns a confidence score of 0 to 1
- Automatic re-routing to a larger model when confidence < 0.85
- Gemma 4 E2B Hybrid matches Gemini 3.1 Flash-Lite performance by routing only 15-35% of queries
- Published benchmark results for ChartQA
- MMBench
- MMLU-Pro
- and others
[DharmaOCR] demonstrated Brazilian Portuguese-specialized OCR model's superiority over newer models [7]
- Superior performance compared to Mistral OCR4 and Unlimited-OCR
- Stage 1 with supervised fine-tuning based on Portuguese files
- Stage 2 stability reinforcement with Direct Preference Optimization
- Achieved the highest extraction quality score and lowest degeneration rate
[NASA Roman Space Telescope] prepared first space-based active coronagraph deployment and exoplanet observation [4]
- Equipped with approximately 300-megapixel wide-field camera
- Expected to detect approximately 100
- 000 new exoplanets
- Implemented active wavefront control using 2 deformable mirrors based on a 48x48 actuator grid
- Approximately 100x wider imaging range compared to Hubble
Sources
- [1] Advancing the next era of national science - OpenAI Blog
- [2] A research agenda for the Economic Futures Research Fund - Anthropic News
- [3] Ask Claude about the Anthropic Economic Index - Anthropic News
- [4] Shape-shifting mirrors on NASA’s new space telescope could unveil Jupiters like our own - MIT Technology Review AI
- [5] Are AI Labs Pelicanmaxxing? - Hacker News
- [6] Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong - Hacker News
- [7] Newer Models, Same Advantage - Hugging Face Blog