collaborators

6 papers

cs.LG2026

Physics-Guided Policy Optimization with Self-Distillation

Ke Wang, Yuning Wu, Haoran Liu +3

Self-distilled policy optimization (SDPO) has become a popular paradigm for LLM post-training, where a model learns from its own predictions conditioned on privileged information.…

cs.CL2026

MOA: Multi-Objective Alignment for Role-Playing Agents

Chonghua Liao, Ke Wang, Yuchuan Wu +3

Role-playing agents (RPAs) require balancing multiple objectives, such as instruction following, persona consistency, and stylistic fidelity, which are not always perfectly aligned…

cs.LG2026

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings

Yuning Wu, Ke Wang, Devin Chen +1

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for post-training reasoning models. However, group-based methods such as Group Relative Po…

cs.CL2025

Agentic Reinforcement Learning with Implicit Step Rewards

Xiaoqian Liu, Ke Wang, Yuchuan Wu +4

Large language models (LLMs) are increasingly developed as autonomous agents using reinforcement learning (agentic RL) that reason and act in interactive environments. However, spa…

cs.CL2025

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning

Xiaoqian Liu, Ke Wang, Yongbin Li +6

Large Language Models (LLMs) have shown impressive reasoning capabilities in well-defined problems with clear solutions, such as mathematics and coding. However, they still struggl…

cs.AI2025

SDPO: Segment-Level Direct Preference Optimization for Social Agents

Aobo Kong, Wentao Ma, Shiwan Zhao +7

Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex social dialogues. Direct Preference Optimization (DPO)…