22 papers
WAT3R: Feedforward Underwater 3D Reconstruction
Jiayi Xu, Jiahao Lu, Ziqiang Zheng +4
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consisten…
SCoPE: Sightline-Coordinate Positional Encoding for Video Diffusion Transformers
Minghao Yin, Jiahao Lu, Wenbo Hu +3
Video diffusion transformers address their tokens by position on the pixel-time grid: an address in the tensor, not in the world. The address we would want, the world point a token…
Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor
Yang Zhao, Jiahao Lu, Bin Huang +2
Narang et al. (2021) evaluated 40+ Transformer modifications at T5-base scale and concluded that most did not transfer. Five years later, the typical working regime has moved to 1-…
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
Chengfeng Zhao, Jiazhi Shu, Yubo Zhao +7
In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior for plausibility and consistenc…
ReFlow: Self-correction Motion Learning for Dynamic Scene Reconstruction
Yanzhe Liang, Ruijie Zhu, Hanzhi Chang +3
We present ReFlow, a unified framework for monocular dynamic scene reconstruction that learns 3D motion in a novel self-correction manner from raw video. Existing methods often suf…
MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE
Ruijie Zhu, Jiahao Lu, Wenbo Hu +4
We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint r…