21 papers
ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision
Delin Mao, Chenghao Sun, Jingwei Song +2
Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different…
Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
Chishui Chen, Yaoyou Fan, Te Sun +11
On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn age…
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
Tongxu Luo, Rongsheng Wang, Jiaxi Bi +22
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional cod…
ICBCBench: An Industry Consortium Benchmark for Financial Deep Research
Weiya Li, Zhiwei Tang, Yizhou He +10
With the rapid advancement of Deep Research Agents in knowledge-intensive domains such as finance, establishing reliable and domain-aligned evaluation standards remains a critical…
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Chuan Xiao, Zhengbo Jiao, Shaobo Wang +5
LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by the availability of high-qualit…
IndustryCode: A Benchmark for Industry Code Generation
Puyu Zeng, Zhaoxi Wang, Zhixu Duan +7
Code generation and comprehension by Large Language Models (LLMs) have emerged as core drivers of industrial intelligence and decision optimization, finding widespread application…