3 papers
cs.AI2026
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
Yongxin Wang, Ruizhe Zhou, Yueling Tang +4
Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text,…
cs.AI2026
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
Kun Xiang, Terry Jingchen Zhang, Zirong Liu +15
We introduce SeePhys Pro, a fine-grained modality transfer benchmark that studies whether models preserve the same reasoning capability when critical information is progressively t…
cs.CL2026
VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios
Yu-Hsiang Liu, Yu-Chien Tang, An-Zi Yen
Evaluating whether AI agents can proactively assist humans in daily activities, ranging from routine household tasks to urgent safety-critical situations, requires diverse visual d…