collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

Jiazhuo Li, Linjiang Cao, Qi Liu +1

Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world models reduce the reliance on c…

cs.LG2026

CoGenCast: A Coupled Autoregressive-Flow Generative Framework for Time Series Forecasting

Mingyue Cheng, Yaguo Liu, Daoyu Wang +2

Time series forecasting can be viewed as a generative problem that requires both semantic understanding over contextual conditions and stochastic modeling of continuous temporal dy…

cs.LG2026

MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning

Xiaoyu Tao, Mingyue Cheng, Ze Guo +4

Time series forecasting (TSF) plays a critical role in decision-making for many real-world applications. Recently, large language model (LLM)- based forecasters have made promising…

cs.LG2026

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning

Daoyu Wang, Mingyue Cheng, Qingchuan Li +3

Agentic reinforcement learning (RL) has become an important post-training paradigm for turning LLMs from static chatbots into interactive agents, giving rise to representative appl…

cs.LG2026

SocraticPO: Policy Optimization via Interactive Guidance

Zirui Liu, Jie Ouyang, Qi Liu +8

Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization dir…

cs.LG2026

Position: Beyond Model-Centric Prediction -- Agentic Time Series Forecasting

Mingyue Cheng, Xiaoyu Tao, Qi Liu +2

Time series forecasting has traditionally been formulated as a model-centric, static, and single-pass prediction problem that maps historical observations to future values. While t…