4 papers
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
Zhiheng Xi, Xin Guo, Yang Nan +18
Reinforcement learning (RL) has recently become the core paradigm for aligning and strengthening large language models (LLMs). Yet, applying RL in off-policy settings--where stale…
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Zhiheng Xi, Jixuan Huang, Chenyang Liao +20
Developing autonomous LLM agents capable of making a series of intelligent decisions to solve complex, real-world tasks is a fast-evolving frontier. Like human cognitive developmen…
Converging/diverging self-similar shock waves: from collapse to reflection
Juhi Jang, Jiaqi Liu, Matthew Schrecker
We solve the continuation problem for the non-isentropic Euler equations following the collapse of an imploding shock wave. More precisely, we prove that the self-similar Güderley…
On self-similar converging shock waves
Juhi Jang, Jiaqi Liu, Matthew Schrecker
In this paper, we rigorously prove the existence of self-similar converging shock wave solutions for the non-isentropic Euler equations for . These solutions are analyt…