5 papers
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
Yanjia Huang, Xianshun Jiang, Xiangbo Gao +2
Vision-and-Language Navigation (VLN) requires agents to follow language instructions while acting in continuous real-world spaces. Prior image imagination based VLN work shows bene…
FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
Yanjia Huang, Shuo Liu, Sheng Liu +4
Long-horizon robot manipulation tasks remain challenging for Vision-Language-Action (VLA) policies due to drift and exposure bias, often denoise the entire trajectory with fixed hy…
VISTA: Generative Visual Imagination for Vision-and-Language Navigation
Yanjia Huang, Mingyang Wu, Renjie Li +1
Vision-and-Language Navigation (VLN) tasks agents with locating specific objects in unseen environments using natural language instructions and visual cues. Many existing VLN appro…
Can Large Vision Language Models Read Maps Like a Human?
Shuo Xing, Zezhou Sun, Shuangyu Xie +6
In this paper, we introduce MapBench-the first dataset specifically designed for human-readable, pixel-based map-based outdoor navigation, curated from complex path finding scenari…
PANDORA: Diffusion Policy Learning for Dexterous Robotic Piano Playing
Yanjia Huang, Renjie Li, Zhengzhong Tu
We present PANDORA, a novel diffusion-based policy learning framework designed specifically for dexterous robotic piano performance. Our approach employs a conditional U-Net archit…