collaborators

5 papers

cs.CV2026

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Renye Yan, Jikang Cheng, Shikun Sun +7

Despite strong image-generation performance, diffusion models' reconstruction objectives limit alignment with human preferences. RL enables such alignment through explicit rewards.…

cs.CV2025

Dual-Flow: Transferable Multi-Target, Instance-Agnostic Attacks via In-the-wild Cascading Flow Optimization

Yixiao Chen, Shikun Sun, Jianshu Li +3

Adversarial attacks are widely used to evaluate model robustness, and in black-box scenarios, the transferability of these attacks becomes crucial. Existing generator-based attacks…

cs.HC2025

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos

Qixin Wang, Songtao Zhou, Zeyu Jin +3

Automatic video commentary systems are widely used on multimedia social media platforms to extract factual information about video content. However, current systems may overlook es…

cs.LG2025

Minimal Impact ControlNet: Advancing Multi-ControlNet Integration

Shikun Sun, Min Zhou, Zixuan Wang +7

With the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple contr…

cs.CV2025

Skinned Motion Retargeting with Dense Geometric Interaction Perception

Zijie Ye, Jia-Wei Liu, Jia Jia +2

Capturing and maintaining geometric interactions among different body parts is crucial for successful motion retargeting in skinned characters. Existing approaches often overlook b…