activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters

Philippe Hansen-Estruch, Jiahui Chen, Vivek Ramanujan +9

Vision Transformer (ViT) autoencoders have emerged as compelling tokenizers for images, offering improved reconstruction over convolutional tokenizers. However, existing ViT tokeni…

cs.CV2026

Unified Text-Image Generation with Weakness-Targeted Post-Training

Jiahui Chen, Philippe Hansen-Estruch, Xiaochuang Han +7

Unified multimodal generation architectures that jointly produce text and images have recently emerged as a promising direction for text-to-image (T2I) synthesis. However, many exi…

cs.CV2025

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Philippe Hansen-Estruch, David Yan, Ching-Yao Chung +7

Visual tokenization via auto-encoding empowers state-of-the-art image and video generative models by compressing pixels into a latent space. Although scaling Transformer-based gene…

cs.CV2024

Apollo: An Exploration of Video Understanding in Large Multimodal Models

Orr Zohar, Xiaohan Wang, Yann Dubois +9

Despite the rapid integration of video perception capabilities into Large Multimodal Models (LMMs), the underlying mechanisms driving their video understanding remain poorly unders…

cs.CV2024

Video Occupancy Models

Manan Tomar, Philippe Hansen-Estruch, Philip Bachman +4

We introduce a new family of video prediction models designed to support downstream control tasks. We call these models Video Occupancy models (VOCs). VOCs operate in a compact lat…

cs.CV2024

Unified Auto-Encoding with Masked Diffusion

Philippe Hansen-Estruch, Sriram Vishwanath, Amy Zhang +1

At the core of both successful generative and self-supervised representation learning models there is a reconstruction objective that incorporates some form of image corruption. Di…