papers

Publications (9)

cs.CL2025

Enough Coin Flips Can Make LLMs Act Bayesian

Ritwik Gupta, Rodolfo Corona, Jiaxin Ge +4

Large language models (LLMs) exhibit the ability to generalize given few-shot examples in their input prompt, an emergent capability known as in-context learning (ICL). We investig…

cs.CL2024

Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding

Chancharik Mitra, Abrar Anwar, Rodolfo Corona +3

When connecting objects and their language referents in an embodied 3D environment, it is important to note that: (1) an object can be better characterized by leveraging comparativ…

cs.CV2024

Analyzing The Language of Visual Tokens

David M. Chan, Rodolfo Corona, Joonyong Park +3

With the introduction of transformer-based models for vision and language tasks, such as LLaVA and Chameleon, there has been renewed interest in the discrete tokenized representati…

cs.AI2019

Modeling Conceptual Understanding in Image Reference Games

Rodolfo Corona, Stephan Alaniz, Zeynep Akata

An agent who interacts with a wide population of other agents needs to be aware that there may be variations in their understanding of the world. Furthermore, the machinery which t…

cs.CL2021

Modular Networks for Compositional Instruction Following

Rodolfo Corona, Daniel Fried, Coline Devin +2

Standard architectures used in instruction following often struggle on novel compositions of subgoals (e.g. navigating to landmarks or picking up objects) observed during training.…

cs.CL2022

Voxel-informed Language Grounding

Rodolfo Corona, Shizhan Zhu, Dan Klein +1

Natural language applied to natural 2D images describes a fundamentally 3D world. We present the Voxel-informed Language Grounder (VLG), a language grounding model that leverages 3…

cs.CV2025

Images are Worth Variable Length of Representations

Lingjun Mao, Rodolfo Corona, Xin Liang +2

Most existing vision encoders map images into a fixed-length sequence of tokens, overlooking the fact that different images contain varying amounts of information. For example, a v…

cs.CL2024

Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction

Boyi Li, Rodolfo Corona, Karttikeya Mangalam +7

Are multimodal inputs necessary for grammar induction? Recent work has shown that multimodal training inputs can improve grammar induction. However, these improvements are based on…

cs.CY2024

Data-Centric AI Governance: Addressing the Limitations of Model-Focused Policies

Ritwik Gupta, Leah Walker, Rodolfo Corona +5

Current regulations on powerful AI capabilities are narrowly focused on "foundation" or "frontier" models. However, these terms are vague and inconsistently defined, leading to an…