3 papers
cs.AI2026
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry
Weiyang Guo, Zesheng Shi, Longhui Zhang +3
Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajector…
cs.CR2026
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
Weiyang Guo, Zesheng Shi, Zeen Zhu +3
Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abilities on complex logical task…
cs.AI2026
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning
Weiyang Guo, Zesheng Shi, Liye Zhao +5
While Large Language Models (LLMs) have demonstrated significant potential in Tool-Integrated Reasoning (TIR), existing training paradigms face significant limitations: Zero-RL suf…