1 citations · 1 across the 3 of their papers we have counts for
5 papers
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
Qinchuan Cheng, Zhantao Gong, Pengzhan Sun +3
Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmark…
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
Zhantao Gong, Liaoyuan Fan, Qing Guo +3
In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applications such as autonomous driving, yet lar…
SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
Peiru Zheng, Yun Zhao, Zhan Gong +2
End-to-end autonomous driving has emerged as a promising paradigm for achieving robust and intelligent driving policies. However, existing end-to-end methods still face significant…
SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection
Yun Zhao, Zhan Gong, Peiru Zheng +2
More and more research works fuse the LiDAR and camera information to improve the 3D object detection of the autonomous driving system. Recently, a simple yet effective fusion fram…
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving
Peiru Zheng, Yun Zhao, Zhan Gong +2
Many fields could benefit from the rapid development of the large language models (LLMs). The end-to-end autonomous driving (e2eAD) is one of the typically fields facing new opport…