6 papers
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
Ruoyu Wang, Jingke Wang, Yukai Ma +5
Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, e…
Reinforcing Few-step Generators via Reward-Tilted Distribution Matching
Yushi Huang, Xiangxin Zhou, Ruoyu Wang +3
Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains challenging. We propose Reward-…
VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding
Ruoyu Wang, Yong Liu, Sheng Tao +2
Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally constructs dense spatial representations on the fly. However, recent Gaussian-centric m…
LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images
Guichen Huang, Ruoyu Wang, Xiangjun Gao +4
3D Gaussian Splatting achieves high-fidelity novel view synthesis, but its application to online long-sequence scenarios is still limited. Existing methods either rely on slow per-…
Recollection from Pensieve: Novel View Synthesis via Learning from Uncalibrated Videos
Ruoyu Wang, Yi Ma, Shenghua Gao
Currently almost all state-of-the-art novel view synthesis and reconstruction models rely on calibrated cameras or additional geometric priors for training. These prerequisites sig…
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
Chun-Hsiao Yeh, Chenyu Wang, Shengbang Tong +7
Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental…