3 papers
cs.CV2026
Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
Yihao Wu, Chenyi Xu, Liqi Yan +6
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multim…
cs.CV2026
World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks
Zuyao Lin, Jianhui Zhang, Peidong Jia +3
World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures…
cs.CV2026
LongDPM: Overlap-Aware 4D Reconstruction from Long Monocular Videos
Chenyi Xu, Yihao Wu, Liqi Yan +4
Recovering a dynamic 3D scene from a long monocular video is crucial for dense geometry, camera motion, and temporal correspondence to remain consistent in a shared coordinate syst…