2 papers
cs.LG2026
TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment
Jiaming Li, Chenyu Zhu, Nanxi Yi +9
Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from significant reward hacking, which…
cs.MM2026
Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD
Junhao Xiao, Shun Feng, Zhiyu Wu +6
Audio-Visual Speaker Detection (AVSD) hinges on modeling both individual temporal continuity and inter-personal social context. Existing coupled architectures struggle to reconcile…