most citedScaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

27 citations · 38 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD2023

FLAP: Fast Language-Audio Pre-training

Ching-Feng Yeh, Po-Yao Huang, Vasu Sharma +2

We propose Fast Language-Audio Pre-training (FLAP), a self-supervised approach that efficiently and effectively learns aligned audio and language representations through masking, c…

cs.LG202327 cited

Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Lili Yu, Bowen Shi, Ramakanth Pasunuru +24

We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. C…

cs.HC20232 cited

Alexa, play with robot: Introducing the First Alexa Prize SimBot Challenge on Embodied AI

Hangjie Shi, Leslie Ball, Govind Thattai +39

The Alexa Prize program has empowered numerous university students to explore, experiment, and showcase their talents in building conversational agents through challenges like the…

cs.HC20233 cited

Alexa Arena: A User-Centric Interactive Platform for Embodied AI

Qiaozi Gao, Govind Thattai, Suhaila Shakiah +24

We introduce Alexa Arena, a user-centric simulation platform for Embodied AI (EAI) research. Alexa Arena provides a variety of multi-room layouts and interactable objects, for the…

cs.AI20226 cited

CH-MARL: A Multimodal Benchmark for Cooperative, Heterogeneous Multi-Agent Reinforcement Learning

Vasu Sharma, Prasoon Goyal, Kaixiang Lin +3

We propose a multimodal (vision-and-language) benchmark for cooperative and heterogeneous multi-agent learning. We introduce a benchmark multimodal dataset with tasks involving col…