4 papers
VC-Tooler: Learning Compositional and Adaptive Visual Tool Use
Yizheng Wu, Jiashen Hua, Bing Deng +1
Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool interactions. Effective visual…
Synthetic-to-Real Translation for Class-Agnostic Motion Prediction
Yizheng Wu, Hongwei Fan, Kewei Wang +8
Motion understanding is critical for ensuring safety and robustness in autonomous driving systems, driving increasing interest in motion prediction. A key challenge in this domain…
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
Zhiyu Pan, Yizheng Wu, Jiashen Hua +5
Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques that refine reasoning paths fo…
Semi-Supervised High Dynamic Range Image Reconstructing via Bi-Level Uncertain Area Masking
Wei Jiang, Jiahao Cui, Yizheng Wu +3
Reconstructing high dynamic range (HDR) images from low dynamic range (LDR) bursts plays an essential role in the computational photography. Impressive progress has been achieved b…