activity
20242026
most citedSimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

Qinchuan Cheng, Zhantao Gong, Pengzhan Sun +3

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmark…

cs.CV2025

Thinking Ahead: Foresight Intelligence in MLLMs and World Model

Zhantao Gong, Liaoyuan Fan, Qing Guo +3

In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applications such as autonomous driving, yet lar…

cs.RO2025

SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving

Peiru Zheng, Yun Zhao, Zhan Gong +2

End-to-end autonomous driving has emerged as a promising paradigm for achieving robust and intelligent driving policies. However, existing end-to-end methods still face significant…

cs.CV20241 cited

SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection

Yun Zhao, Zhan Gong, Peiru Zheng +2

More and more research works fuse the LiDAR and camera information to improve the 3D object detection of the autonomous driving system. Recently, a simple yet effective fusion fram…

cs.CV2024

SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving

Peiru Zheng, Yun Zhao, Zhan Gong +2

Many fields could benefit from the rapid development of the large language models (LLMs). The end-to-end autonomous driving (e2eAD) is one of the typically fields facing new opport…