From the 1 of 13 linked papers with an AI index.
13 papers
Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models
Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang +5
The paper presents a large-scale structured reasoning dataset created via slice‑wise synthesis that encodes chain‑of‑thought explanations for 3D medical images, and uses it to inst…
DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
Hangui Lin, Yan Shu, Zhengyang Liang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a leading paradigm for enhancing visual reasoning in Multimodal Large Language Models (MLLMs). However, existin…
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
Yue Zhang, Zun Wang, Han Lin +3
Spatial reasoning is a fundamental capability for vision-language models (VLMs) deployed in real-world environments. However, visual observations are inherently limited representat…
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
Zun Wang, Jaemin Cho, Jialu Li +4
Recent approaches for video generation with camera control often create anchor videos (i.e., rendered videos that approximate desired camera motions) to guide diffusion models as a…
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
Yidong Huang, Zun Wang, Han Lin +6
Generating realistic human motion is a central yet unsolved challenge in video generation. While reinforcement learning (RL)-based post-training has driven recent gains in general…
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
Han Lin, Xichen Pan, Zun Wang +4
Pixel-space diffusion has recently re-emerged as a strong alternative to latent diffusion, enabling high-quality generation without pretrained autoencoders. However, standard pixel…