3 papers
cs.CV2026
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
Jianglin Lu, Simon Jenni, Kushal Kafle +3
Text-to-image retrieval is a fundamental task in vision-language learning, yet in real-world scenarios it is often challenged by short and underspecified user queries. Such queries…
cs.AI2026
Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
Guanyu Yao, Qiucheng Wu, Yang Zhang +3
Multimodal large language models (MLLMs) have demonstrated strong capabilities on vision-and-language tasks. However, recent findings reveal an imbalance in their reasoning capabil…
cs.CV2025
VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos
Qiucheng Wu, Handong Zhao, Zhixin Shu +3
Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they sti…