Publications (11)
Pix2seq: A Language Modeling Framework for Object Detection
Ting Chen, Saurabh Saxena, Lala Li +2
We present Pix2Seq, a simple and generic framework for object detection. Unlike existing approaches that explicitly integrate prior knowledge about the task, we cast object detecti…
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia, William Chan, Saurabh Saxena +11
We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large tran…
Big Bidirectional Insertion Representations for Documents
Lala Li, William Chan
The Insertion Transformer is well suited for long form text generation due to its parallel generation capabilities, requiring generation steps to generate tokens.…
Controlling Space and Time with Diffusion Models
Daniel Watson, Saurabh Saxena, Lala Li +2
We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, condition…
Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
Guodong Zhang, Lala Li, Zachary Nado +5
Increasing the batch size is a popular way to speed up neural network training, but beyond some critical batch size, larger batch sizes yield diminishing returns. In this work, we…
A Generalist Framework for Panoptic Segmentation of Images and Videos
Ting Chen, Lala Li, Saurabh Saxena +2
Panoptic segmentation assigns semantic and instance ID labels to every pixel of an image. As permutations of instance IDs are also valid solutions, the task requires learning of hi…