activity
20242026
most citedLearn A Flexible Exploration Model for Parameterized Action Markov Decision Processes

1 citations · 1 across the 11 of their papers we have counts for

collaborators

18 papers

cs.LG2026

Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes

Amrith Setlur, Zijian Wang, Andrew Cohen +2

Typical reinforcement learning (RL) methods for LLM reasoning waste compute on hard problems, where correct on-policy traces are rare, policy gradients vanish, and learning stalls.…

cs.LG2026

PaperGuide: Making Small Language-Model Paper-Reading Agents More Efficient

Zijian Wang, Tiancheng Huang, Hanqi Li +3

The accelerating growth of the scientific literature makes it increasingly difficult for researchers to track new advances through manual reading alone. Recent progress in large la…

cs.AI2026

ThinkDrive: Chain-of-Thought Guided Progressive Reinforcement Learning Fine-Tuning for Autonomous Driving

Chang Zhao, Zheming Yang, Yunqing Hu +4

With the rapid advancement of large language models (LLMs) technologies, their application in the domain of autonomous driving has become increasingly widespread. However, existing…

cs.RO2025

Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation

Siyu Xu, Zijian Wang, Yunke Wang +3

Vision-Language-Action (VLA) models have shown great performance in robotic manipulation by mapping visual observations and language instructions directly to actions. However, they…

cs.CL2025

Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization

Zijian Wang, Yanxiang Ma, Chang Xu

Chain-of-Thought (CoT) reasoning is a critical capability for large language models (LLMs), enabling them to tackle com- plex multi-step tasks. While base LLMs, pre-trained on gene…

cs.LG2025

Controllable Flow Matching for Online Reinforcement Learning

Bin Wang, Boxiang Tao, Haifeng Jing +2

Model-based reinforcement learning (MBRL) typically relies on modeling environment dynamics for data efficiency. However, due to the accumulation of model errors over long-horizon…