Google I/O 2026 Showcases Multimodal AI Expansion Centered on Gemini Omni
멀티모달 | Wed Jun 17 2026 00:00:00 GMT+0000 (Coordinated Universal Time) | 14 sources
From the unveiling of Gemini Omni and 3.5 Flash to smart glasses and world models, multimodal AI expanded across the industry.
Analysis
[Google] unveiled Gemini Omni and Gemini 3.5 Flash [1][4][5]
- Generates high-quality video from image
- audio
- video
- and text inputs
- Enables video editing through natural language conversation
- Gemini 3.5 Flash delivers frontier performance specialized for agent and coding tasks
- Rolling out across Gemini app
- Google Flow
- and YouTube Shorts
[Google] introduced multimodal Search inputs and an information agent [5]
- Uses text
- images
- files
- video
- and Chrome tabs as search inputs
- Launched an information agent that operates 24/7 in the background
- Creates customized agents via 'keep me updated' command
- Available first to Google AI Pro and Ultra subscribers
[Google] produced I/O 2026 video and brand assets using Gemini [2][3]
- Generated stylized first frames of raw footage with Nano Banana
- Combined puppet animation with AI frames using Gemini Omni
- Maintained frame consistency with custom tools in Google AI Studio
- Non-developers also created Antigravity-based vibe coding quizzes
[Meta] expanded Facebook AI Mode and Ray-Ban Meta glasses [6][7][8]
- Muse Spark-based AI Mode generates answers from public posts
- AI photo presets change outfits
- hair
- and accessories
- Provided Ray-Ban Meta free to 130
- 000 blind U.S. veterans
- Supports object recognition and text reading by voice
[Qualcomm, Snap, and Apple] accelerated competition in multimodal hardware for smart glasses [9][10][11]
- Qualcomm Snapdragon Reality Elite boosts NPU performance by up to 160%
- Supports 4.4K 90fps binocular and 20% improved battery
- Snap Specs AR glasses revealed at around $2
- 200
- sending stock down 5%
- Apple exploring camera-equipped AirPods to give Siri visual context
[Odyssey, Allen AI, and Pinterest] expanded into world models, 3D motion, and visual search [12][13][14]
- Odyssey reached $1.45B valuation in Series B with Amazon and AMD participation
- Provides world model optimized for AWS Trainium chips
- Allen AI MolmoMotion predicts 3D point trajectories from language instructions
- Pinterest experimentally launched conversational shopping app 'Ask Pinterest'
Sources
- [1] The latest AI news we announced in May 2026 - Google AI Blog
- [2] How we used Gemini to build Google I/O 2026 - Google AI Blog
- [3] Take our I/O 2026 quiz, vibe coded in Google AI Studio. - Google AI Blog
- [4] 9 demos of Gemini Omni and Gemini 3.5 in action - Google AI Blog
- [5] Catch up on 12 major I/O 2026 moments - Google AI Blog
- [6] New AI Tools to Help You Make Things Happen on Facebook - Meta AI Blog
- [7] The Future Is for Everyone: Free AI Glasses for Every Blind Veteran in America - Meta AI Blog
- [8] AI search grounded in Facebook posts? What could go wrong? - The Verge AI
- [9] Apple 2027 rumors: AirPods with cameras for AI and the second folding iPhone - The Verge AI
- [10] Qualcomm’s latest chip hints that more powerful smart glasses could be on the way - The Verge AI
- [11] After unveiling ridiculously expensive AR glasses, Snap’s stock takes a dive - TechCrunch AI
- [12] World model maker Odyssey nabs $1.45B valuation backed by Amazon and other big names - TechCrunch AI
- [13] Pinterest launches an experimental AI shopping app called ‘Ask Pinterest’ - TechCrunch AI
- [14] MolmoMotion: Language-guided 3D motion forecasting - Hugging Face Blog