2 papers
cs.CV2026
Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs
Chongyu Wang, Ting Huang, Chunyu Sun +3
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in 2D visual tasks but still struggle to understand physical space in real-world visual streams. Recently…
cs.CV2023
R3D-SWIN:Use Shifted Window Attention for Single-View 3D Reconstruction
Chenhuan Li, Meihua Xiao, zehuan li +5
Recently, vision transformers have performed well in various computer vision tasks, including voxel 3D reconstruction. However, the windows of the vision transformer are not multi-…