7 papers · 1 filter
From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
Jiajia Song, Bobo Li, Haiwen Yi +6
Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have be…
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry
Weiyang Guo, Zesheng Shi, Longhui Zhang +3
Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajector…
AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning
Yutong Wang, Siyuan Xiong, Xuebo Liu +4
While Multi-Agent Systems (MAS) excel in complex reasoning, they suffer from the cascading impact of erroneous information from individual agents. Current solutions often resort to…
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning
Xinyu Ma, Mingzhou Xu, Xuebo Liu +4
Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet models often struggle to explore…
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
Benteng Chen, Weida Wang, Shufei Zhang +2
Large reasoning models that use long chain-of-thought excel at problem-solving yet waste compute on redundant checks. Curbing this overthinking is hard: training-time length penalt…
Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
Jun Rao, Xuebo Liu, Hexuan Deng +5
In mathematical reasoning, data selection strategies predominantly rely on static, externally defined metrics, which fail to adapt to the evolving capabilities of models during tra…