273 citations · 518 across the 26 of their papers we have counts for
8 papers · 1 filter
Diffusion Illusions: Hiding Images in Plain Sight
Ryan Burgert, Xiang Li, Abe Leite +2
We explore the problem of computationally generating special `prime' images that produce optical illusions when physically arranged and viewed in a certain way. First, we propose a…
Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities
AJ Piergiovanni, Isaac Noble, Dahun Kim +3
One of the main challenges of multimodal learning is the need to combine heterogeneous modalities (e.g., video, audio, text). For example, video and audio are obtained at much high…
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Anthony Brohan, Noah Brown, Justice Carbajal +51
We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic…
Language-based Action Concept Spaces Improve Video Self-Supervised Learning
Kanchana Ranasinghe, Michael Ryoo
Recent contrastive language image pre-training has led to learning highly transferable and robust image representations. However, adapting these models to video domains with minima…
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
Xiang Li, Varun Belagali, Jinghuan Shang +1
Sequence modeling approaches have shown promising results in robot imitation learning. Recently, diffusion models have been adopted for behavioral cloning in a sequence modeling fa…
Energy-Based Models for Cross-Modal Localization using Convolutional Transformers
Alan Wu, Michael S. Ryoo
We present a novel framework using Energy-Based Models (EBMs) for localizing a ground vehicle mounted with a range sensor against satellite imagery in the absence of GPS. Lidar sen…