activity
20232026
collaborators

7 papers

cs.AI2026

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

Shudong Liu, Dongyang Chen, Enci Zhang +3

Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexte…

cs.CL2026

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Changhao Xiang, Shilin Zhang, Zheng Ma +8

Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capabili…

cs.AI2026

OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

Kanzhi Cheng, Zehao Li, Zheng Ma +11

Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achieving a marked performance leap…

cs.CV2025

CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Kanzhi Cheng, Wenpo Song, Jiaxin Fan +7

Image captioning has been a longstanding challenge in vision-language research. With the rise of LLMs, modern Vision-Language Models (VLMs) generate detailed and comprehensive imag…

cs.AI2024

MixRED: A Mix-lingual Relation Extraction Dataset

Lingxing Kong, Yougang Chu, Zheng Ma +3

Relation extraction is a critical task in the field of natural language processing with numerous real-world applications. Existing research primarily focuses on monolingual relatio…

cs.MM2024

Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description

Mianzhi Pan, Jianfei Li, Mingyue Yu +4

Commonsense reasoning, the ability to make logical assumptions about daily scenes, is one core intelligence of human beings. In this work, we present a novel task and dataset for e…