6 papers
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning
Shijin Gong, Kai Ye, Jin Zhu +3
Recent advances in large language models (LLMs) have increasingly relied on reinforcement learning (RL) to improve their reasoning capabilities. Three types of approaches have been…
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
V Team, Wenyi Hong, Xiaotao Gu +94
We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depen…
What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators
Xinyu Zhang
World models learn to simulate environment dynamics from experience, enabling sample-efficient reinforcement learning. But what do these models actually represent internally? We ap…
Reliable Self-Improvement Training by Verifying Reasoning, Not Just Answers
Xinyu Zhang
Self-improvement training, where models learn from self-generated solutions, promises sustained capability gains but suffers from a pervasive failure mode: across multiple rounds,…
: Discovering Logical Formulaic Alphas using Deep Reinforcement Learning
Feng Xu, Yan Yin, Xinyu Zhang +3
Alphas are pivotal in providing signals for quantitative trading. The industry highly values the discovery of formulaic alphas for their interpretability and ease of analysis, comp…
Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics
Xinyu Zhang, Wenjie Qiu, Yi-Chen Li +4
Developing policies that can adjust to non-stationary environments is essential for real-world reinforcement learning applications. However, learning such adaptable policies in off…