11 papers
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
Kishan Panaganti, Zhenwen Liang, Wenhao Yu +2
Recent progress in Large Language Model (LLM) reasoning is increasingly driven by the refinement of post-training loss functions and alignment strategies. However, standard Reinfor…
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
Zhenwen Liang, Sidi Lu, Wenhao Yu +4
Reinforcement learning has become essential for strengthening the reasoning abilities of large language models, yet current exploration mechanisms remain fundamentally misaligned w…
Guided Self-Evolving LLMs with Minimal Human Supervision
Wenhao Yu, Zhenwen Liang, Chengsong Huang +4
AI self-evolution has long been envisioned as a path toward superintelligence, where models autonomously acquire, refine, and internalize knowledge from their own learning experien…
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
Dian Yu, Yulai Zhao, Kishan Panaganti +3
We propose Reinforcement Learning with Explicit Human Values (RLEV), a method that aligns Large Language Model (LLM) optimization directly with quantifiable human value signals. Wh…
TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
Héctor Carrión, Yutong Bai, Víctor A. Hernández Castro +6
World models aim to simulate environments and enable effective agent behavior. However, modeling real-world environments presents unique challenges as they dynamically change acros…
Risk-Averse Total-Reward Reinforcement Learning
Xihong Su, Jia Lin Hau, Gersi Doko +2
Risk-averse total-reward Markov Decision Processes (MDPs) offer a promising framework for modeling and solving undiscounted infinite-horizon objectives. Existing model-based algori…