130 citations · 587 across the 58 of their papers we have counts for
11 papers · 1 filter
Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text
Christopher Clark, Jordi Salvador, Dustin Schwenk +13
Communicating with humans is challenging for AIs because it requires a shared understanding of the world, complex semantics (e.g., metaphors or analogies), and at times multi-modal…
Asymptotic behavior of the free interface for entire vector minimizers in phase transitions
Nicholas D. Alikakos, Zhiyuan Geng, Arghir Zarnescu
We study globally bounded entire minimizers of Allen-Cahn systems for potentials with and $W(u)\sim |u-a…
Simple but Effective: CLIP Embeddings for Embodied AI
Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi +1
Contrastive language image pretraining (CLIP) encoders have been shown to be beneficial for a range of visual tasks from classification and detection to captioning and image manipu…
RobustNav: Towards Benchmarking Robustness in Embodied Navigation
Prithvijit Chattopadhyay, Judy Hoffman, Roozbeh Mottaghi +1
As an attempt towards assessing the robustness of embodied navigation agents, we propose RobustNav, a framework to quantify the performance of embodied navigation agents when expos…
Container: Context Aggregation Network
Peng Gao, Jiasen Lu, Hongsheng Li +2
Convolutional neural networks (CNNs) are ubiquitous in computer vision, with a myriad of effective and efficient variations. Recently, Transformers -- originally introduced in natu…
PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D World
Rowan Zellers, Ari Holtzman, Matthew Peters +4
We propose PIGLeT: a model that learns physical commonsense knowledge through interaction, and then uses this knowledge to ground language. We factorize PIGLeT into a physical dyna…