4 papers
The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
Xinyu Jessica Wang, Haoyue Bai, Yiyou Sun +7
Large language model (LLM) agents perform strongly on short- and mid-horizon tasks, but often break down on long-horizon tasks that require extended, interdependent action sequence…
RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
Jianxing Liao, Tian Zhang, Xiao Feng +6
Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emot…
-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
Peijie Yu, Yifan Yang, Jinjian Li +4
Agents based on large language models leverage tools to modify environments, revolutionizing how AI interacts with the physical world. Unlike traditional NLP tasks that rely solely…
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
Peijie Yu, Yifan Yang, Jinjian Li +4
Large language models (LLMs) demonstrate strong potential as agents for tool invocation due to their advanced comprehension and planning capabilities. Users increasingly rely on LL…