3 papers
cs.LG2026
Offline Reinforcement Learning with Generative Trajectory Policies
Xinsong Feng, Leshu Tang, Chenan Wang +1
Generative models have emerged as a powerful class of policies for offline reinforcement learning (RL) due to their ability to capture complex, multi-modal behaviors. However, exis…
cs.LG2026
Speculative Sampling with Reinforcement Learning
Chenan Wang, Daniel H. Shi, Haipeng Chen
Inference time latency has remained an open challenge for real world applications of large language models (LLMs). State-of-the-art (SOTA) speculative sampling (SpS) methods for LL…
cs.CL2026
DIP: Dynamic In-Context Planner For Diffusion Language Models
Yang Li, Han Meng, Chenan Wang +1
Diffusion language models (DLMs) have shown strong potential for general natural language tasks with in-context examples. However, due to the bidirectional attention mechanism, DLM…