activity
20242026
collaborators

5 papers

cs.CV2026

MMGist: A Comprehensive Multimodal Benchmark for 2027

Wenzhen Yuan, Jiacheng Ruan, Wutao Xiong +3

We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and therefore fail to effective…

cs.AI2026

Small Model as Master Orchestrator: Learning Unified Agent-Tool Orchestration with Parallel Subtask Decomposition

Wenzhen Yuan, Wutao Xiong, Fanchen Yu +7

Multi-agent systems (MAS) demonstrate clear advantages in tackling complex problems by coordinating diverse agents and external tools. However, most existing orchestration methods…

cs.AI2025

Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback

Wenzhen Yuan, Shengji Tang, Weihao Lin +8

Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), but its reliance on expensive human-labeled data or complex rewar…

cs.CV2025

VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models

Jiacheng Ruan, Wenzhen Yuan, Xian Gao +6

Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Re…

cs.CV2024

MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object Scenarios

Jiacheng Ruan, Wenzhen Yuan, Zehao Lin +5

Large visual-language models (LVLMs) have achieved great success in multiple applications. However, they still encounter challenges in complex scenes, especially those involving ca…