Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
Chao Yu, Qixin Tan, Jiaxuan Gao +7
Reasoning reinforcement learning (RL) has recently revealed a new scaling effect: test-time scaling. Thinking models such as R1 and o1 improve their reasoning accuracy at test time…
cs.LG2025
Fine-tuning Diffusion Policies with Backpropagation Through Diffusion Timesteps
Ningyuan Yang, Jiaxuan Gao, Feng Gao +2
Diffusion policies, widely adopted in decision-making scenarios such as robotics, gaming and autonomous driving, are capable of learning diverse skills from demonstration data due…