12 papers
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Haozhe Wang, Weijia Feng, Jinpeng Yu +8
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending…
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
Yuhuan Wu, Cong Wei, Fangzhen Lin +2
Vision-Language Models (VLMs) deployed as situated agents in high-resolution visual environments require active perception -- the ability to dynamically decide where to look throug…
Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning
Haozhe Wang, Qixin Xu, Changpeng Wang +4
Achieving robust perception-reasoning synergy is a central goal for advanced Vision-Language Models (VLMs). Recent advancements have pursued this goal via architectural designs or…
Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
Tianyi Xiong, Yi Ge, Ming Li +13
Large multimodal models (LMMs) are increasingly adopted as judges in multimodal evaluation systems due to their strong instruction following and consistency with human preferences.…
Towards Trustworthy GUI Agents: A Survey
Yucheng Shi, Wenhao Yu, Jingyuan Huang +3
Graphical User Interface (GUI) agents extend large language models from text generation to action execution in real-world digital environments. Unlike conversational systems, GUI a…
Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
Chenlong Wang, Yuhang Chen, Zhihan Hu +4
Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are gen…