4 papers · 1 filter
ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning
Xiaoxuan Wang, Han Zhang, Haixin Wang +11
Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks. Despite encouraging ea…
HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness
Xiaoxuan Wang, Haixin Wang, Alexander Taylor +3
Large language models are increasingly deployed as agents for long-horizon tasks, yet their performance is shaped not only by model capability and environment design, but also by t…
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
Junkai Zhang, Jingru Gan, Xiaoxuan Wang +8
Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied. To fill this gap, we introduce MatSc…
Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
Lianhao Zhou, Hongyi Ling, Cong Fu +14
Computing has long served as a cornerstone of scientific discovery. Recently, a paradigm shift has emerged with the rise of large language models (LLMs), introducing autonomous sys…