7 papers
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
Madhuri Shanbhogue, Zhe Li, Shanfeng Zhang +86
We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage t…
Learning to Learn from Language Feedback with Social Meta-Learning
Jonathan Cook, Diego Antognini, Martin Klissarov +2
Large language models (LLMs) often struggle to learn from corrective feedback within a conversational context. They are rarely proactive in soliciting this feedback, even when face…
Improving Interactive In-Context Learning from Natural Language Feedback
Martin Klissarov, Jonathan Cook, Diego Antognini +5
Adapting one's thought process based on corrective feedback is an essential ability in human learning, particularly in collaborative settings. In contrast, the current large langua…
Position: Introspective Experience from Conversational Environments as a Path to Better Learning
Claudiu Cristian Musat, Jackson Tolins, Diego Antognini +3
Current approaches to AI training treat reasoning as an emergent property of scale. We argue instead that robust reasoning emerges from linguistic self-reflection, itself internali…
Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
Riccardo Brioschi, Aleksandr Alekseev, Emanuele Nevali +9
Graphic layout generation is a growing research area focusing on generating aesthetically pleasing layouts ranging from poster designs to documents. While recent research has explo…
InkSight: Offline-to-Online Handwriting Conversion by Teaching Vision-Language Models to Read and Write
Blagoj Mitrevski, Arina Rak, Julian Schnitzler +4
Digital note-taking is gaining popularity, offering a durable, editable, and easily indexable way of storing notes in a vectorized form, known as digital ink. However, a substantia…