33 citations · 33 across the 5 of their papers we have counts for
4 papers · 1 filter
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
Andrew Szot, Bogdan Mazoure, Omar Attia +6
We examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typica…
On the Modeling Capabilities of Large Language Models for Sequential Decision Making
Martin Klissarov, Devon Hjelm, Alexander Toshev +1
Large pretrained models are showing increasingly better performance in reasoning and planning tasks across different modalities, opening the possibility to leverage them for comple…
Grounding Multimodal Large Language Models in Actions
Andrew Szot, Bogdan Mazoure, Harsh Agrawal +3
Multimodal Large Language Models (MLLMs) have demonstrated a wide range of capabilities across many domains, including Embodied AI. In this work, we study how to best ground a MLLM…
Poly-View Contrastive Learning
Amitis Shidani, Devon Hjelm, Jason Ramapuram +3
Contrastive learning typically matches pairs of related views among a number of unrelated negative views. Views can be generated (e.g. by augmentations) or be observed. We investig…