7 papers
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
Shudong Liu, Dongyang Chen, Enci Zhang +3
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexte…
OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories
Changhao Xiang, Shilin Zhang, Zheng Ma +8
Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capabili…
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
Kanzhi Cheng, Zehao Li, Zheng Ma +11
Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achieving a marked performance leap…
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
Kanzhi Cheng, Wenpo Song, Jiaxin Fan +7
Image captioning has been a longstanding challenge in vision-language research. With the rise of LLMs, modern Vision-Language Models (VLMs) generate detailed and comprehensive imag…
MixRED: A Mix-lingual Relation Extraction Dataset
Lingxing Kong, Yougang Chu, Zheng Ma +3
Relation extraction is a critical task in the field of natural language processing with numerous real-world applications. Existing research primarily focuses on monolingual relatio…
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
Mianzhi Pan, Jianfei Li, Mingyue Yu +4
Commonsense reasoning, the ability to make logical assumptions about daily scenes, is one core intelligence of human beings. In this work, we present a novel task and dataset for e…