12 papers
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Bo-Wen Zhang, Junwei He, Wen Wang +5
The paper introduces CoRT, a method that uses counterfactual replay to assign token-level credit in rubric-guided reinforcement learning for language models, improving credit alloc…
CAST: Game Solvers as Turn-Level Teachers for LLM Agents
Yu Wang, Yi-Kai Zhang, Wentao Shi +8
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR)…
Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
Hong-Jie You, Jie-Jing Shao, Xiao-Wen Yang +3
Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance from a symbolic score, rely on supervised…
Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models
Xiao-Wen Yang, Ziyu Han, Xi-Hua Zhang +4
Looped Language Models (LoopLMs) enable efficient latent reasoning through depth recurrence, yet exhibit unreliable test-time scaling behavior: performance often peaks at a certain…
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
Jiaxuan Wang, Yulan Hu, Wenjin Yang +3
In classical Reinforcement Learning from Human Feedback (RLHF), Reward Models (RMs) serve as the fundamental signal provider for model alignment. As Large Language Models evolve in…
Revisiting the Travel Planning Capabilities of Large Language Models
Bo-Wen Zhang, Jin Ye, Peng-Yu Hua +4
Travel planning serves as a critical task for long-horizon reasoning, exposing significant deficits in LLMs. However, existing benchmarks and evaluations primarily assess final pla…