2 papers
cs.CL2026
EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports
Yu-Chien Tang, Yu-Hsiang Liu, An-Zi Yen
VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly ev…
cs.AI2026
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
Yongxin Wang, Ruizhe Zhou, Yueling Tang +4
Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text,…