Krea 2 Image Model Release and FFASR Benchmark Launch
모델 출시 | Thu Jun 25 2026 00:00:00 GMT+0000 (Coordinated Universal Time) | 2 sources
Krea released a 12B open-weights image generation model, while Hugging Face and Treble launched a far-field ASR benchmark.
Analysis
[Krea] released Krea 2, a 12B open-weights image generation model [1]
- SOTA open-weights model at 12B parameters
- Based on diffusion transformer (DiT) architecture
- Foundation model optimized for creative exploration
[Krea 2] applied a multi-stage training pipeline and architectural improvements [1]
- Stages of pretraining
- midtraining
- SFT
- preference optimization
- and RL
- Integration of iREPA
- improved VAE
- and Qwen3-VL
- Introduction of grouped-query attention (GQA) and sigmoid-gated attention
- Lightweight timestep modulation and multilayer feature aggregation
[Krea 2] introduced a prompt expander and style-reference system [1]
- Expands short
- ambiguous user prompts into rich visual descriptions
- Steerable from both text and image inputs
- Narrows the gap between training conditioning and user expressions at inference
[Hugging Face · Treble Technologies] launched the FFASR Leaderboard [2]
- The first open far-field ASR benchmark
- Community-driven evaluation based on 14 simulated rooms
- Real-world measurements and validated sim-to-real methodology
- Pareto front visualization of WER and RTFx
[FFASR Benchmark] confirmed a large performance gap compared to near-field [2]
- Far-field WER is several times higher than near-field in low-SNR environments
- Reflects real-world use cases such as AI voice agents
- conference room transcription
- and in-vehicle assistants
- Roadmap includes support for multi-talker
- microphone arrays
- and echo cancellation