5 papers · 1 filter
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints
Jiayu Liu, Cheng Qian, Zhenhailong Wang +10
Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are progressively disclosed through…
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
Hyeonjeong Ha, Jeonghwan Kim, Cheng Qian +7
Memory-augmented large language models extend reasoning beyond a fixed context window by maintaining long-term memory across interactions. However, existing memory systems often co…
UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind
Cheng Qian, Jiayu Liu, Heng Ji
Understanding what a user believes and intends is central to building effective agent assistants. This ability is often evaluated through Theory-of-Mind (ToM) tasks, where success…
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
Dadi Guo, Yuejin Xie, Qingyu Liu +11
As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality problems has become a significa…
Diversity-Enhanced Reasoning for Subjective Questions
Yumeng Wang, Zhiyuan Fan, Jiayu Liu +2
Large Reasoning Models (LRMs) with long chain-of-thought capabilities, optimized via reinforcement learning with verifiable rewards (RLVR), excel at objective reasoning tasks like…