3 papers
cs.LG2026
Step-level Denoising-time Diffusion Alignment with Multiple Objectives
Qi Zhang, Dawei Wang, Shaofeng Zou
Reinforcement learning (RL) has emerged as a powerful tool for aligning diffusion models with human preferences, typically by optimizing a single reward function under a KL regular…
cs.LG2026
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
Keru Chen, Jun Luo, Sen Lin +4
Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO…
cs.LG2025
LDC-MTL: Balancing Multi-Task Learning through Scalable Loss Discrepancy Control
Peiyao Xiao, Chaosheng Dong, Shaofeng Zou +1
Multi-task learning (MTL) has been widely adopted for its ability to simultaneously learn multiple tasks. While existing gradient manipulation methods often yield more balanced sol…