27 citations · 38 across the 5 of their papers we have counts for
5 papers
FLAP: Fast Language-Audio Pre-training
Ching-Feng Yeh, Po-Yao Huang, Vasu Sharma +2
We propose Fast Language-Audio Pre-training (FLAP), a self-supervised approach that efficiently and effectively learns aligned audio and language representations through masking, c…
Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
Lili Yu, Bowen Shi, Ramakanth Pasunuru +24
We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. C…
Alexa, play with robot: Introducing the First Alexa Prize SimBot Challenge on Embodied AI
Hangjie Shi, Leslie Ball, Govind Thattai +39
The Alexa Prize program has empowered numerous university students to explore, experiment, and showcase their talents in building conversational agents through challenges like the…
Alexa Arena: A User-Centric Interactive Platform for Embodied AI
Qiaozi Gao, Govind Thattai, Suhaila Shakiah +24
We introduce Alexa Arena, a user-centric simulation platform for Embodied AI (EAI) research. Alexa Arena provides a variety of multi-room layouts and interactable objects, for the…
CH-MARL: A Multimodal Benchmark for Cooperative, Heterogeneous Multi-Agent Reinforcement Learning
Vasu Sharma, Prasoon Goyal, Kaixiang Lin +3
We propose a multimodal (vision-and-language) benchmark for cooperative and heterogeneous multi-agent learning. We introduce a benchmark multimodal dataset with tasks involving col…