activity
20242026
collaborators

8 papers

cs.RO2026

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

Mengfei Zhao, Dihong Huang, Yikai Tang +12

Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on speciali…

cs.CV2026

Fishbone: From One 3D Asset to a Million Controllable Edits

Yumeng He, Xiaoying Wang, Peihao Li +7

Large-scale controllable 3D assets are critical for computer graphics, embodied AI, robotics, and interactive content creation, yet creating diverse 3D assets remains challenging d…

cs.RO2025

VISTAv2: World Imagination for Indoor Vision-and-Language Navigation

Yanjia Huang, Xianshun Jiang, Xiangbo Gao +2

Vision-and-Language Navigation (VLN) requires agents to follow language instructions while acting in continuous real-world spaces. Prior image imagination based VLN work shows bene…

cs.RO2025

FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation

Yanjia Huang, Shuo Liu, Sheng Liu +4

Long-horizon robot manipulation tasks remain challenging for Vision-Language-Action (VLA) policies due to drift and exposure bias, often denoise the entire trajectory with fixed hy…

cs.RO2025

VISTA: Generative Visual Imagination for Vision-and-Language Navigation

Yanjia Huang, Mingyang Wu, Renjie Li +1

Vision-and-Language Navigation (VLN) tasks agents with locating specific objects in unseen environments using natural language instructions and visual cues. Many existing VLN appro…

cs.CV2025

Can Large Vision Language Models Read Maps Like a Human?

Shuo Xing, Zezhou Sun, Shuangyu Xie +6

In this paper, we introduce MapBench-the first dataset specifically designed for human-readable, pixel-based map-based outdoor navigation, curated from complex path finding scenari…