51 citations · 87 across the 15 of their papers we have counts for
12 papers · 1 filter
Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation
Niange Yu, Ye Tian, Biaolong Chen +5
Multi-subject video generation faces two key challenges: uncontrollable fidelity strength and potential semantic drift. We address these by analyzing the internal mechanisms of Dif…
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
Shiyi Zhang, Mushui Liu, Yunze Tong +8
On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory signals, has achieved significant success in Large Language Mode…
RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation
Yuhan Li, Xianfeng Tan, Fangao Zeng +6
Standard clothing asset generation---restoring forward-facing flat-lay garment images from diverse real-world contexts---holds immense commercial value yet demands both macroscopic…
InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
Yunze Tong, Mushui Liu, Canyu Zhao +9
With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align t…
RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation
Yuhan Li, Fangao Zeng, Sicong Kang +5
Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequenti…
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Peng Zhang, Guanghao Zhang, Wanggui He +10
Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video…