4 citations · 4 across the 3 of their papers we have counts for
2 papers
cs.CV2026
On-Policy Self-Distillation in Diffusion Models
Wei Zhou, Xiongwei Zhu, Lingdong Kong +14
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction…
cs.CV2022★ 4 cited
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval
Mengjun Cheng, Yipeng Sun, Longchao Wang +8
Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable…