5 papers · 1 filter
Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains
Yanchao Li, Wanhao Liu, Jiaqing Xie +4
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produc…
SkillsInjector: Dynamic Skill Context Construction for LLM Agents
Yanchao Li, Wanhao Liu, Ben Gao +5
LLM agents now draw on growing skill libraries to handle complex tasks. However, injecting more skills does not always improve task completion and can even degrade it. Existing met…
MolAct: An Agentic RL Framework for Molecular Editing and Property Optimization
Zhuo Yang, Yeyun Chen, Jiaqing Xie +7
Molecular editing and optimization are multi-step problems that require iteratively improving properties while keeping molecules chemically valid and structurally similar. We frame…
DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
Haiyuan Wan, Chen Yang, Junchi Yu +10
Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and e…
QCBench: Evaluating Large Language Models on Domain-Specific Quantitative Chemistry
Jiaqing Xie, Weida Wang, Ben Gao +5
Quantitative chemistry is central to modern chemical research, yet the ability of large language models (LLMs) to perform its rigorous, step-by-step calculations remains underexplo…