8 citations · 8 across the 6 of their papers we have counts for
4 papers · 1 filter
Hearing Anywhere in Any Environment
Xiulong Liu, Anurag Kumar, Paul Calamia +7
In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances…
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta +4
Leveraging Large Language Models' remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. Howeve…
The ObjectFolder Benchmark: Multisensory Learning with Neural and Real Objects
Ruohan Gao, Yiming Dou, Hao Li +5
We introduce the ObjectFolder Benchmark, a benchmark suite of 10 tasks for multisensory object-centric learning, centered around object recognition, reconstruction, and manipulatio…
An Extensible Multimodal Multi-task Object Dataset with Materials
Trevor Standley, Ruohan Gao, Dawn Chen +2
We present EMMa, an Extensible, Multimodal dataset of Amazon product listings that contains rich Material annotations. It contains more than 2.8 million objects, each with image(s)…