5 papers
Trust Region Policy Distillation
Zhengpeng Xie, Li Lyna Zhang, Zeke Xie +1
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable,…
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
Lichen Bai, Tianhao Zhang, Shitong Shao +14
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but…
CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Think
Zening Sun, Zhengpeng Xie, Lichen Bai +3
Aligning Diffusion models has achieved remarkable breakthroughs in generating high-quality, human preference-aligned images. Existing techniques, such as supervised fine-tuning (SF…
Zeroth-Order Optimization is Secretly Single-Step Policy Optimization
Junbin Qiu, Zhengpeng Xie, Xiangda Yan +2
Zeroth-Order Optimization (ZOO) provides powerful tools for optimizing functions where explicit gradients are unavailable or expensive to compute. However, the underlying mechanism…
Simple Policy Optimization
Zhengpeng Xie, Qiang Zhang, Fan Yang +2
Model-free reinforcement learning algorithms have seen remarkable progress, but key challenges remain. Trust Region Policy Optimization (TRPO) is known for ensuring monotonic polic…