Artwork

āđ€āļ™āļ·āđ‰āļ­āļŦāļēāļˆāļąāļ”āļ—āļģāđ‚āļ”āļĒ TWIML and Sam Charrington āđ€āļ™āļ·āđ‰āļ­āļŦāļēāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāļ—āļąāđ‰āļ‡āļŦāļĄāļ” āļĢāļ§āļĄāļ–āļķāļ‡āļ•āļ­āļ™ āļāļĢāļēāļŸāļīāļ āđāļĨāļ°āļ„āļģāļ­āļ˜āļīāļšāļēāļĒāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāđ„āļ”āđ‰āļĢāļąāļšāļāļēāļĢāļ­āļąāļ›āđ‚āļŦāļĨāļ”āđāļĨāļ°āļˆāļąāļ”āļŦāļēāđƒāļŦāđ‰āđ‚āļ”āļĒāļ•āļĢāļ‡āļˆāļēāļ TWIML and Sam Charrington āļŦāļĢāļ·āļ­āļžāļąāļ™āļ˜āļĄāļīāļ•āļĢāđāļžāļĨāļ•āļŸāļ­āļĢāđŒāļĄāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāļ‚āļ­āļ‡āļžāļ§āļāđ€āļ‚āļē āļŦāļēāļāļ„āļļāļ“āđ€āļŠāļ·āđˆāļ­āļ§āđˆāļēāļĄāļĩāļšāļļāļ„āļ„āļĨāļ­āļ·āđˆāļ™āđƒāļŠāđ‰āļ‡āļēāļ™āļ—āļĩāđˆāļĄāļĩāļĨāļīāļ‚āļŠāļīāļ—āļ˜āļīāđŒāļ‚āļ­āļ‡āļ„āļļāļ“āđ‚āļ”āļĒāđ„āļĄāđˆāđ„āļ”āđ‰āļĢāļąāļšāļ­āļ™āļļāļāļēāļ• āļ„āļļāļ“āļŠāļēāļĄāļēāļĢāļ–āļ›āļāļīāļšāļąāļ•āļīāļ•āļēāļĄāļ‚āļąāđ‰āļ™āļ•āļ­āļ™āļ—āļĩāđˆāđāļŠāļ”āļ‡āđ„āļ§āđ‰āļ—āļĩāđˆāļ™āļĩāđˆ https://th.player.fm/legal
Player FM - āđāļ­āļ› Podcast
āļ­āļ­āļŸāđ„āļĨāļ™āđŒāļ”āđ‰āļ§āļĒāđāļ­āļ› Player FM !

Inside Nano Banana 🍌 and the Future of Vision-Language Models with Oliver Wang - #748

1:03:39
 
āđāļšāđˆāļ‡āļ›āļąāļ™
 

Manage episode 508093774 series 2355587
āđ€āļ™āļ·āđ‰āļ­āļŦāļēāļˆāļąāļ”āļ—āļģāđ‚āļ”āļĒ TWIML and Sam Charrington āđ€āļ™āļ·āđ‰āļ­āļŦāļēāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāļ—āļąāđ‰āļ‡āļŦāļĄāļ” āļĢāļ§āļĄāļ–āļķāļ‡āļ•āļ­āļ™ āļāļĢāļēāļŸāļīāļ āđāļĨāļ°āļ„āļģāļ­āļ˜āļīāļšāļēāļĒāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāđ„āļ”āđ‰āļĢāļąāļšāļāļēāļĢāļ­āļąāļ›āđ‚āļŦāļĨāļ”āđāļĨāļ°āļˆāļąāļ”āļŦāļēāđƒāļŦāđ‰āđ‚āļ”āļĒāļ•āļĢāļ‡āļˆāļēāļ TWIML and Sam Charrington āļŦāļĢāļ·āļ­āļžāļąāļ™āļ˜āļĄāļīāļ•āļĢāđāļžāļĨāļ•āļŸāļ­āļĢāđŒāļĄāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāļ‚āļ­āļ‡āļžāļ§āļāđ€āļ‚āļē āļŦāļēāļāļ„āļļāļ“āđ€āļŠāļ·āđˆāļ­āļ§āđˆāļēāļĄāļĩāļšāļļāļ„āļ„āļĨāļ­āļ·āđˆāļ™āđƒāļŠāđ‰āļ‡āļēāļ™āļ—āļĩāđˆāļĄāļĩāļĨāļīāļ‚āļŠāļīāļ—āļ˜āļīāđŒāļ‚āļ­āļ‡āļ„āļļāļ“āđ‚āļ”āļĒāđ„āļĄāđˆāđ„āļ”āđ‰āļĢāļąāļšāļ­āļ™āļļāļāļēāļ• āļ„āļļāļ“āļŠāļēāļĄāļēāļĢāļ–āļ›āļāļīāļšāļąāļ•āļīāļ•āļēāļĄāļ‚āļąāđ‰āļ™āļ•āļ­āļ™āļ—āļĩāđˆāđāļŠāļ”āļ‡āđ„āļ§āđ‰āļ—āļĩāđˆāļ™āļĩāđˆ https://th.player.fm/legal

Today, we’re joined by Oliver Wang, principal scientist at Google DeepMind and tech lead for Gemini 2.5 Flash Image—better known by its code name, “Nano Banana.” We dive into the development and capabilities of this newly released frontier vision-language model, beginning with the broader shift from specialized image generators to general-purpose multimodal agents that can use both visual and textual data for a variety of tasks. Oliver explains how Nano Banana can generate and iteratively edit images while maintaining consistency, and how its integration with Gemini’s world knowledge expands creative and practical use cases. We discuss the tension between aesthetics and accuracy, the relative maturity of image models compared to text-based LLMs, and scaling as a driver of progress. Oliver also shares surprising emergent behaviors, the challenges of evaluating vision-language models, and the risks of training on AI-generated data. Finally, we look ahead to interactive world models and VLMs that may one day “think” and “reason” in images.

The complete show notes for this episode can be found at https://twimlai.com/go/748.

  continue reading

768 āļ•āļ­āļ™

Artwork
iconāđāļšāđˆāļ‡āļ›āļąāļ™
 
Manage episode 508093774 series 2355587
āđ€āļ™āļ·āđ‰āļ­āļŦāļēāļˆāļąāļ”āļ—āļģāđ‚āļ”āļĒ TWIML and Sam Charrington āđ€āļ™āļ·āđ‰āļ­āļŦāļēāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāļ—āļąāđ‰āļ‡āļŦāļĄāļ” āļĢāļ§āļĄāļ–āļķāļ‡āļ•āļ­āļ™ āļāļĢāļēāļŸāļīāļ āđāļĨāļ°āļ„āļģāļ­āļ˜āļīāļšāļēāļĒāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāđ„āļ”āđ‰āļĢāļąāļšāļāļēāļĢāļ­āļąāļ›āđ‚āļŦāļĨāļ”āđāļĨāļ°āļˆāļąāļ”āļŦāļēāđƒāļŦāđ‰āđ‚āļ”āļĒāļ•āļĢāļ‡āļˆāļēāļ TWIML and Sam Charrington āļŦāļĢāļ·āļ­āļžāļąāļ™āļ˜āļĄāļīāļ•āļĢāđāļžāļĨāļ•āļŸāļ­āļĢāđŒāļĄāļžāļ­āļ”āđāļ„āļŠāļ•āđŒāļ‚āļ­āļ‡āļžāļ§āļāđ€āļ‚āļē āļŦāļēāļāļ„āļļāļ“āđ€āļŠāļ·āđˆāļ­āļ§āđˆāļēāļĄāļĩāļšāļļāļ„āļ„āļĨāļ­āļ·āđˆāļ™āđƒāļŠāđ‰āļ‡āļēāļ™āļ—āļĩāđˆāļĄāļĩāļĨāļīāļ‚āļŠāļīāļ—āļ˜āļīāđŒāļ‚āļ­āļ‡āļ„āļļāļ“āđ‚āļ”āļĒāđ„āļĄāđˆāđ„āļ”āđ‰āļĢāļąāļšāļ­āļ™āļļāļāļēāļ• āļ„āļļāļ“āļŠāļēāļĄāļēāļĢāļ–āļ›āļāļīāļšāļąāļ•āļīāļ•āļēāļĄāļ‚āļąāđ‰āļ™āļ•āļ­āļ™āļ—āļĩāđˆāđāļŠāļ”āļ‡āđ„āļ§āđ‰āļ—āļĩāđˆāļ™āļĩāđˆ https://th.player.fm/legal

Today, we’re joined by Oliver Wang, principal scientist at Google DeepMind and tech lead for Gemini 2.5 Flash Image—better known by its code name, “Nano Banana.” We dive into the development and capabilities of this newly released frontier vision-language model, beginning with the broader shift from specialized image generators to general-purpose multimodal agents that can use both visual and textual data for a variety of tasks. Oliver explains how Nano Banana can generate and iteratively edit images while maintaining consistency, and how its integration with Gemini’s world knowledge expands creative and practical use cases. We discuss the tension between aesthetics and accuracy, the relative maturity of image models compared to text-based LLMs, and scaling as a driver of progress. Oliver also shares surprising emergent behaviors, the challenges of evaluating vision-language models, and the risks of training on AI-generated data. Finally, we look ahead to interactive world models and VLMs that may one day “think” and “reason” in images.

The complete show notes for this episode can be found at https://twimlai.com/go/748.

  continue reading

768 āļ•āļ­āļ™

āļ—āļļāļāļ•āļ­āļ™

×
 
Loading …

āļ‚āļ­āļ•āđ‰āļ­āļ™āļĢāļąāļšāļŠāļđāđˆ Player FM!

Player FM āļāļģāļĨāļąāļ‡āļŦāļēāđ€āļ§āđ‡āļš

 

āļ„āļđāđˆāļĄāļ·āļ­āļ­āđ‰āļēāļ‡āļ­āļīāļ‡āļ”āđˆāļ§āļ™

āļŸāļąāļ‡āļĢāļēāļĒāļāļēāļĢāļ™āļĩāđ‰āđƒāļ™āļ‚āļ“āļ°āļ—āļĩāđˆāļ„āļļāļ“āļŠāļģāļĢāļ§āļˆ
āđ€āļĨāđˆāļ™