5 papers
Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation
Wanrong Zheng, Yunhao Ge, Laurent Itti
Breakthrough progress in vision-based navigation through unknown environments has been achieved by using multimodal large language models (MLLMs). These models can plan a sequence…
Value Explicit Pretraining for Learning Transferable Representations
Kiran Lekkala, Henghui Bao, Sumedh A. Sontakke +2
Understanding visual inputs for a given task amidst varied changes is a key challenge posed by visual reinforcement learning agents. We propose \textit{Value Explicit Pretraining}…
Towards Embodiment Scaling Laws in Robot Locomotion
Bo Ai, Liu Dai, Nico Bohlinger +7
Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…
DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation
Brian Nlong Zhao, Yuhang Xiao, Jiashu Xu +6
The popularization of Text-to-Image (T2I) diffusion models enables the generation of high-quality images from text descriptions. However, generating diverse customized images with…
Perforated Backpropagation: A Neuroscience Inspired Extension to Artificial Neural Networks
Rorry Brenner, Laurent Itti
The neurons of artificial neural networks were originally invented when much less was known about biological neurons than is known today. Our work explores a modification to the co…