papers

Publications (13)

cs.CV2024

HAMMR: HierArchical MultiModal React agents for generic VQA

Lluis Castrejon, Thomas Mensink, Howard Zhou +3

Combining Large Language Models (LLMs) with external specialized tools (LLMs+tools) is a recent paradigm to solve multimodal tasks such as Visual Question Answering (VQA). While th…

cs.CV2023

How (not) to ensemble LVLMs for VQA

Lisa Alazraki, Lluis Castrejon, Mostafa Dehghani +3

This paper studies ensembling in the era of Large Vision-Language Models (LVLMs). Ensembling is a classical method to combine different models to get increased performance. In the…

cs.CV2019

Improved Conditional VRNNs for Video Prediction

Lluis Castrejon, Nicolas Ballas, Aaron Courville

Predicting future frames for a video sequence is a challenging generative modeling task. Promising approaches include probabilistic latent variable models such as the Variational A…

cs.CV2023

Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories

Thomas Mensink, Jasper Uijlings, Lluis Castrejon +6

We propose Encyclopedic-VQA, a large scale visual question answering (VQA) dataset featuring visual questions about detailed properties of fine-grained categories and instances. It…

cs.CV2016

Cross-Modal Scene Networks

Yusuf Aytar, Lluis Castrejon, Carl Vondrick +2

People can recognize scenes across many different modalities beyond natural images. In this paper, we investigate how to learn cross-modal scene representations that transfer acros…

cs.CV2022

Cascaded Video Generation for Videos In-the-Wild

Lluis Castrejon, Nicolas Ballas, Aaron Courville

Videos can be created by first outlining a global view of the scene and then adding local details. Inspired by this idea we propose a cascaded model for video generation which foll…

cs.LG2020

Supervision Accelerates Pre-training in Contrastive Semi-Supervised Learning of Visual Representations

Mahmoud Assran, Nicolas Ballas, Lluis Castrejon +1

We investigate a strategy for improving the efficiency of contrastive learning of visual representations by leveraging a small amount of supervised information during pre-training.…

cs.CV2021

Hierarchical Video Generation for Complex Data

Lluis Castrejon, Nicolas Ballas, Aaron Courville

Videos can often be created by first outlining a global description of the scene and then adding local details. Inspired by this we propose a hierarchical model for video generatio…

cs.CV2018

MovieGraphs: Towards Understanding Human-Centric Situations from Videos

Paul Vicol, Makarand Tapaswi, Lluis Castrejon +1

There is growing interest in artificial intelligence to build socially intelligent robots. This requires machines to have the ability to "read" people's emotions, motivations, and…

cs.CV2024

Imagen 3

Imagen-Team-Google, :, Jason Baldridge +257

We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred…

cs.CV2017

Annotating Object Instances with a Polygon-RNN

Lluis Castrejon, Kaustav Kundu, Raquel Urtasun +1

We propose an approach for semi-automatic annotation of object instances. While most current methods treat object segmentation as a pixel-labeling problem, we here cast it as a pol…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

cs.CV2016

Learning Aligned Cross-Modal Representations from Weakly Aligned Data

Lluis Castrejon, Yusuf Aytar, Carl Vondrick +2

People can recognize scenes across many different modalities beyond natural images. In this paper, we investigate how to learn cross-modal scene representations that transfer acros…