4 papers
Trust Your Guide Only When Certain: Uncertainty-Aware Sparse Alignment at Inference Time
Zeen Zhu, Zhuo Li, Weiyang Guo +4
A prominent paradigm in inference-time alignment employs lightweight supervisors to steer Large Language Models (LLMs). Through empirical analysis, we identify a structural mismatc…
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry
Weiyang Guo, Zesheng Shi, Longhui Zhang +3
Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajector…
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
Weiyang Guo, Zesheng Shi, Zeen Zhu +3
Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abilities on complex logical task…
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning
Weiyang Guo, Zesheng Shi, Liye Zhao +5
While Large Language Models (LLMs) have demonstrated significant potential in Tool-Integrated Reasoning (TIR), existing training paradigms face significant limitations: Zero-RL suf…