Publications (9)
Enough Coin Flips Can Make LLMs Act Bayesian
Ritwik Gupta, Rodolfo Corona, Jiaxin Ge +4
Large language models (LLMs) exhibit the ability to generalize given few-shot examples in their input prompt, an emergent capability known as in-context learning (ICL). We investig…
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
Chancharik Mitra, Abrar Anwar, Rodolfo Corona +3
When connecting objects and their language referents in an embodied 3D environment, it is important to note that: (1) an object can be better characterized by leveraging comparativ…
Analyzing The Language of Visual Tokens
David M. Chan, Rodolfo Corona, Joonyong Park +3
With the introduction of transformer-based models for vision and language tasks, such as LLaVA and Chameleon, there has been renewed interest in the discrete tokenized representati…
Modeling Conceptual Understanding in Image Reference Games
Rodolfo Corona, Stephan Alaniz, Zeynep Akata
An agent who interacts with a wide population of other agents needs to be aware that there may be variations in their understanding of the world. Furthermore, the machinery which t…
Modular Networks for Compositional Instruction Following
Rodolfo Corona, Daniel Fried, Coline Devin +2
Standard architectures used in instruction following often struggle on novel compositions of subgoals (e.g. navigating to landmarks or picking up objects) observed during training.…
Voxel-informed Language Grounding
Rodolfo Corona, Shizhan Zhu, Dan Klein +1
Natural language applied to natural 2D images describes a fundamentally 3D world. We present the Voxel-informed Language Grounder (VLG), a language grounding model that leverages 3…
Images are Worth Variable Length of Representations
Lingjun Mao, Rodolfo Corona, Xin Liang +2
Most existing vision encoders map images into a fixed-length sequence of tokens, overlooking the fact that different images contain varying amounts of information. For example, a v…
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction
Boyi Li, Rodolfo Corona, Karttikeya Mangalam +7
Are multimodal inputs necessary for grammar induction? Recent work has shown that multimodal training inputs can improve grammar induction. However, these improvements are based on…
Data-Centric AI Governance: Addressing the Limitations of Model-Focused Policies
Ritwik Gupta, Leah Walker, Rodolfo Corona +5
Current regulations on powerful AI capabilities are narrowly focused on "foundation" or "frontier" models. However, these terms are vague and inconsistently defined, leading to an…