3 papers
cs.CV2026
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
Houcheng Jiang, Jiajun Fu, Junfeng Fang +4
Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on explicit textual chain-of-tho…
cs.CV2026
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
Shida Gao, Feng Xue, Xiangfeng Wang +8
Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-temporal video grounding (STVG) and re…
cs.CV2025
NTIRE 2025 Challenge on HR Depth from Images of Specular and Transparent Surfaces
Pierluigi Zama Ramirez, Fabio Tosi, Luigi Di Stefano +36
This paper reports on the NTIRE 2025 challenge on HR Depth From images of Specular and Transparent surfaces, held in conjunction with the New Trends in Image Restoration and Enhanc…