4 papers
Drifting Preference Optimization for One-Step Generative Models
Zhou Jiang, Yandong Wen, Zhen Liu
One-step text-to-image generators are attractive for deployment because they generate an image with a single forward pass, but preference finetuning them remains difficult: standar…
Rethinking Preference Alignment for Diffusion Models with Classifier-Free Guidance
Zhou Jiang, Yandong Wen, Zhen Liu
Aligning large-scale text-to-image diffusion models with nuanced human preferences remains challenging. While direct preference optimization (DPO) is simple and effective, large-sc…
Meta-Black-Box-Optimization through Offline Q-function Learning
Zeyuan Ma, Zhiguang Cao, Zhou Jiang +2
Recent progress in Meta-Black-Box-Optimization (MetaBBO) has demonstrated that using RL to learn a meta-level policy for dynamic algorithm configuration (DAC) over an optimization…
RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
Weixin Mao, Weiheng Zhong, Zhou Jiang +8
Existing robot policies predominantly adopt the task-centric approach, requiring end-to-end task data collection. This results in limited generalization to new tasks and difficulti…