depth completion 1dynamic attention 1surface normal estimation 1transformer models 1video geometry estimation 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Towards Consistent Video Geometry Estimation
Zhu Yu, Jingnan Gao, Runmin Zhang +9
ViGeo is a transformer-based model that estimates dense, temporally consistent geometry (depth, surface normals, and point maps) from video sequences using dynamic chunking attenti…
cs.CV2026
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation
Yuxuan Yao, Yuxuan Chen, Hui Li +6
Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visua…
cs.CV2024
MVImgNet2.0: A Larger-scale Dataset of Multi-view Images
Xiaoguang Han, Yushuang Wu, Luyue Shi +7
MVImgNet is a large-scale dataset that contains multi-view images of ~220k real-world objects in 238 classes. As a counterpart of ImageNet, it introduces 3D visual signals via mult…