From the 1 of 17 linked papers with an AI index.
17 papers
Native Video-Action Pretraining for Generalizable Robot Control
Qihang Zhang, Lin Li, Luyao Zhang +26
The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…
Next Forcing: Causal World Modeling with Multi-Chunk Prediction
Gangwei Xu, Qihang Zhang, Jiaming Zhou +4
Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow training convergence and limited co…
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
Cheng Chi, Xianqi Wang, Hongcheng Luo +9
High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their pe…
PCSTracker: Long-Term Scene Flow Estimation for Point Cloud Sequences
Min Lin, Gangwei Xu, Xianqi Wang +2
Point cloud scene flow estimation is fundamental to long-term and fine-grained 3D motion analysis. However, existing methods are typically limited to pairwise settings and struggle…
PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
Xianqi Wang, Hao Yang, Hangtian Wang +4
Modern stereo matching methods have leveraged monocular depth foundation models to achieve superior zero-shot generalization performance. However, most existing methods primarily f…
Pixel-Perfect Visual Geometry Estimation
Gangwei Xu, Haotong Lin, Hongcheng Luo +6
Recovering clean and accurate geometry from images is essential for robotics and augmented reality. However, existing geometry foundation models still suffer severely from flying p…