activity
20242026
collaborators

23 papers

cs.CV2026

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

Ying Li, Xiaobao Wei, Jiajun Cao +10

World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video o…

cs.RO2026

IOI: Decoupling Kinematics and Physics for Interactive World Models

Chengyu Bai, Peidong Jia, Tiecheng Guo +11

Developing generalist embodied agents requires interactive environments providing visually realistic feedback and accurate action-conditioned dynamics. Interactive world models add…

cs.RO2026

PIGEON: VLM-Driven Object Navigation via Points of Interest Selection

Cheng Peng, Zhenzhe Zhang, Xiaobao Wei +7

Object navigation in unseen indoor environments requires agents to perform semantic search under partial observability. Vision-language models (VLMs) provide strong semantic-spatia…

cs.RO2026

GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors

Dongli Wu, Xiaobao Wei, Hao Wang +5

Robotic grasping is a fundamental capability in robotic manipulation. Yet grasping remains challenging under partial observations. Reliable grasping depends on both local contact c…

cs.CV2026

SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene Simulation

Qingpo Wuwu, Xiaobao Wei, Peng Chen +6

While 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives to capture fine details, lea…

cs.CV2026

Feed-Forward Gaussian Splatting from Sparse Aerial Views

Dongli Wu, Zhuoxiao Li, Tongyan Hua +4

Reconstructing large-scale urban scenes from sparse aerial views is a crucial yet challenging task. Due to biased top-down and shallow-oblique camera poses, sparse aerial captures…