11 papers
Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
Tianshuo Xu, Yichen Xie, Depu Meng +5
Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity p…
LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation
Bo Jiang, Depu Meng, Yihan Hu +3
Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedie…
EndoVGGT: GNN-Enhanced Depth Estimation for Surgical 3D Reconstruction
Falong Fan, Yi Xie, Arnis Lektauers +2
Accurate 3D reconstruction of deformable soft tissues is essential for surgical robotic perception. However, low-texture surfaces, specular highlights, and instrument occlusions of…
Multi-Camera View Scaling for Data-Efficient Robot Imitation Learning
Yichen Xie, Yixiao Wang, Shuqi Zhao +4
The generalization ability of imitation learning policies for robotic manipulation is fundamentally constrained by the diversity of expert demonstrations, while collecting demonstr…
RT-GS: Gaussian Splatting with Reflection and Transmittance Primitives
Kunnong Zeng, Chensheng Peng, Yichen Xie +2
Gaussian Splatting is a powerful tool for reconstructing diffuse scenes, but it struggles to simultaneously model specular reflections and the appearance of objects behind semi-tra…
UniQueR: Unified Query-based Feedforward 3D Reconstruction
Chensheng Peng, Quentin Herau, Jiezhi Yang +6
We present UniQueR, a unified query-based feedforward framework for efficient and accurate 3D reconstruction from unposed images. Existing feedforward models such as DUSt3R, VGGT,…