19 papers
MEMO: Multimodal Evidence Memory Organization for Long-Horizon LLM Agents
Xian Gao, Jinpeng Wang, Jiacheng Ruan +3
Long-running LLM agents rely on external memory to store and reuse information beyond a single context window, yet there is a fundamental tension between the continuous accumulatio…
Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
Zongyun Zhang, Jiacheng Ruan, Xian Gao +5
Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code pres…
Contrastive On-Policy Distillation
Jiacheng Ruan, Jun Tang, Wenzhen Yuan +5
On-policy Distillation (OPD) supervises a student model on trajectories sampled from its own policy by minimizing the divergence between the output distributions of the teacher and…
MMGist: A Comprehensive Multimodal Benchmark for 2027
Wenzhen Yuan, Jiacheng Ruan, Wutao Xiong +3
We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and therefore fail to effective…
BasketHAR: A Multimodal Dataset for Human Activity Recognition and Sport Analysis in Basketball Training Scenarios
Xian Gao, Haoyue Zhang, Zongyun Zhang +3
Human Activity Recognition (HAR) involves the automatic identification of user activities and has gained significant research interest due to its broad applicability. Most HAR syst…
Small Model as Master Orchestrator: Learning Unified Agent-Tool Orchestration with Parallel Subtask Decomposition
Wenzhen Yuan, Wutao Xiong, Fanchen Yu +7
Multi-agent systems (MAS) demonstrate clear advantages in tackling complex problems by coordinating diverse agents and external tools. However, most existing orchestration methods…