5 papers
LocusGS: Spatially Grounded Tokens for Feed-Forward 3D Gaussian Splatting
Wenyu Li, Sidun Liu, Tongrui Hu +2
Recent query-based feed-forward 3DGS methods represent a scene using learnable queries, each aggregating multi-view evidence and decoding a group of Gaussians. Ideally, different q…
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
Wenjie Hu, Sidun Liu, Peng Qiao +2
Recent advances in Transformer-based Neural Operators have enabled significant progress in data-driven solvers for Partial Differential Equations (PDEs). Most current research has…
Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
Wenyu Li, Sidun Liu, Peng Qiao +2
We present Muskie, a native multi-view vision backbone designed for 3D vision tasks. Unlike existing models, which are frame-wise and exhibit limited multi-view consistency, Muskie…
Regist3R: Incremental Registration with Stereo Foundation Model
Sidun Liu, Wenyu Li, Peng Qiao +1
Multi-view 3D reconstruction has remained an essential yet challenging problem in the field of computer vision. While DUSt3R and its successors have achieved breakthroughs in 3D re…
Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction
Wenyu Li, Sidun Liu, Peng Qiao +1
Recent advances in data-driven geometric multi-view 3D reconstruction foundation models (e.g., DUSt3R) have shown remarkable performance across various 3D vision tasks, facilitated…