722 citations · 2k across the 48 of their papers we have counts for
13 papers · 1 filter
Multimodal Neurons in Pretrained Text-Only Transformers
Sarah Schwettmann, Neil Chowdhury, Samuel Klein +2
Language models demonstrate remarkable capacity to generalize representations learned in one modality to downstream tasks in other modalities. Can we trace this ability to individu…
Follow Anything: Open-set detection, tracking, and following in real-time
Alaa Maalouf, Ninad Jadhav, Krishna Murthy Jatavallabhula +5
Tracking and following objects of interest is critical to several robotics use cases, ranging from industrial automation to logistics and warehousing, to healthcare and security. I…
DreamTeacher: Pretraining Image Backbones with Deep Generative Models
Daiqing Li, Huan Ling, Amlan Kar +5
In this work, we introduce a self-supervised feature representation learning framework DreamTeacher that utilizes generative networks for pre-training downstream image backbones. W…
Background Prompting for Improved Object Depth
Manel Baradad, Yuanzhen Li, Forrester Cole +4
Estimating the depth of objects from a single image is a valuable task for many vision, robotics, and graphics applications. However, current methods often fail to produce accurate…
Unsupervised Compositional Concepts Discovery with Text-to-Image Generative Models
Nan Liu, Yilun Du, Shuang Li +2
Text-to-image generative models have enabled high-resolution image synthesis across different domains, but require users to specify the content they wish to generate. In this paper…
Improving Factuality and Reasoning in Language Models through Multiagent Debate
Yilun Du, Shuang Li, Antonio Torralba +2
Large language models (LLMs) have demonstrated remarkable capabilities in language generation, understanding, and few-shot learning in recent years. An extensive body of work has e…