Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
DiPOD: Diffusion Policy Optimization without Drifting Apart
Haozhe Jiang, Haiwen Feng, Pieter Abbeel +3
RL post-training has become increasingly pivotal for improving diffusion policies, but existing diffusion policy-gradient methods are often unstable and cannot achieve reliable pol…
cs.LG2026
LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning
Hao Jiang, Enneng Yang, Guojie Zhu +7
Continual learning capability is critical for Industrial LLMs, as deployed models must be continuously updated to meet evolving requirements and environments, rather than repeatedl…