5 papers
MemCtrl: Using MLLMs as Active Memory Controllers on Embodied Agents
Vishnu Sashank Dorbala, Dinesh Manocha
Foundation models rely on in-context learning for personalized decision making. The limited size of this context window necessitates memory compression and retrieval systems like R…
Sensitivity to Redirected Walking Considering Gaze, Posture, and Luminance
Niall L. Williams, Logan C. Stevens, Aniket Bera +1
We study the correlations between redirected walking (RDW) rotation gains and patterns in users' posture and gaze data during locomotion in virtual reality (VR). To do this, we con…
Is the House Ready For Sleeptime? Generating and Evaluating Situational Queries for Embodied Question Answering
Vishnu Sashank Dorbala, Prasoon Goyal, Robinson Piramuthu +3
We present and tackle the problem of Embodied Question Answering (EQA) with Situational Queries (S-EQA) in a household environment. Unlike prior EQA work tackling simple queries th…
Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
Uttaran Bhattacharya, Aniket Bera, Dinesh Manocha
We present a multimodal learning-based method to simultaneously synthesize co-speech facial expressions and upper-body gestures for digital characters using RGB video data captured…
Improving Zero-Shot ObjectNav with Generative Communication
Vishnu Sashank Dorbala, Vishnu Dutt Sharma, Pratap Tokekar +1
We propose a new method for improving zero-shot ObjectNav that aims to utilize potentially available environmental percepts for navigational assistance. Our approach takes into acc…