2 papers
cs.RO2026
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
Yucheng Hu, Jianke Zhang, Yuanfei Luo +9
Equipping embodied agents with the ability to reason about tasks, foresee physical outcomes, and generate precise actions is essential for general-purpose manipulation. While recen…
cs.RO2026
Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning
Zixuan Wang, Huang Fang, Shaoan Wang +4
While large vision-language models (VLMs) show promise for object goal navigation, current methods still struggle with low success rates and inefficient localization of unseen obje…