5 papers · 1 filter
ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training
Zhihao Zhang, Mingqi Wu, Qiaole Dong +15
Continual post-training of large multimodal models should add new capabilities while preserving those from pre-training, and the two goals pull in opposite directions. SFT gives ex…
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
Jiaqiang Li, Yajie Yang, Zhiheng Xi +15
Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-s…
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
Zhiheng Xi, Dingwen Yang, Jiaqi Liu +21
Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic eval…
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
Yuhui Wang, Zhixiong Yang, Ming Zhang +10
In the maintenance of complex systems, fault trees are used to locate problems and provide targeted solutions. To enable fault trees stored as images to be directly processed by la…
MagicAgent: Towards Generalized Agent Planning
Xuhui Ren, Shaokang Dong, Chen Yang +21
The evolution of Large Language Models (LLMs) from passive text processors to autonomous agents has established planning as a core component of modern intelligence. However, achiev…