2 papers
cs.LG2026
Adaptive Replay Buffer for Offline-to-Online Reinforcement Learning
Chihyeon Song, Jaewoo Lee, Jinkyoo Park
Offline-to-Online Reinforcement Learning (O2O RL) faces a critical dilemma in balancing the use of a fixed offline dataset with newly collected online experiences. Standard methods…
cs.LG2026
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
Hyeongyu Kang, Jaewoo Lee, Woocheol Shin +2
Diffusion models excel at generating high-likelihood samples but often require alignment with downstream objectives. Existing fine-tuning methods for diffusion models significantly…