Publications (13)
HAMMR: HierArchical MultiModal React agents for generic VQA
Lluis Castrejon, Thomas Mensink, Howard Zhou +3
Combining Large Language Models (LLMs) with external specialized tools (LLMs+tools) is a recent paradigm to solve multimodal tasks such as Visual Question Answering (VQA). While th…
How (not) to ensemble LVLMs for VQA
Lisa Alazraki, Lluis Castrejon, Mostafa Dehghani +3
This paper studies ensembling in the era of Large Vision-Language Models (LVLMs). Ensembling is a classical method to combine different models to get increased performance. In the…
Improved Conditional VRNNs for Video Prediction
Lluis Castrejon, Nicolas Ballas, Aaron Courville
Predicting future frames for a video sequence is a challenging generative modeling task. Promising approaches include probabilistic latent variable models such as the Variational A…
Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
Thomas Mensink, Jasper Uijlings, Lluis Castrejon +6
We propose Encyclopedic-VQA, a large scale visual question answering (VQA) dataset featuring visual questions about detailed properties of fine-grained categories and instances. It…
Cross-Modal Scene Networks
Yusuf Aytar, Lluis Castrejon, Carl Vondrick +2
People can recognize scenes across many different modalities beyond natural images. In this paper, we investigate how to learn cross-modal scene representations that transfer acros…
Cascaded Video Generation for Videos In-the-Wild
Lluis Castrejon, Nicolas Ballas, Aaron Courville
Videos can be created by first outlining a global view of the scene and then adding local details. Inspired by this idea we propose a cascaded model for video generation which foll…
Supervision Accelerates Pre-training in Contrastive Semi-Supervised Learning of Visual Representations
Mahmoud Assran, Nicolas Ballas, Lluis Castrejon +1
We investigate a strategy for improving the efficiency of contrastive learning of visual representations by leveraging a small amount of supervised information during pre-training.…
Hierarchical Video Generation for Complex Data
Lluis Castrejon, Nicolas Ballas, Aaron Courville
Videos can often be created by first outlining a global description of the scene and then adding local details. Inspired by this we propose a hierarchical model for video generatio…
MovieGraphs: Towards Understanding Human-Centric Situations from Videos
Paul Vicol, Makarand Tapaswi, Lluis Castrejon +1
There is growing interest in artificial intelligence to build socially intelligent robots. This requires machines to have the ability to "read" people's emotions, motivations, and…
Imagen 3
Imagen-Team-Google, :, Jason Baldridge +257
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred…
Annotating Object Instances with a Polygon-RNN
Lluis Castrejon, Kaustav Kundu, Raquel Urtasun +1
We propose an approach for semi-automatic annotation of object instances. While most current methods treat object segmentation as a pixel-labeling problem, we here cast it as a pol…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Learning Aligned Cross-Modal Representations from Weakly Aligned Data
Lluis Castrejon, Yusuf Aytar, Carl Vondrick +2
People can recognize scenes across many different modalities beyond natural images. In this paper, we investigate how to learn cross-modal scene representations that transfer acros…