activity
20182026
most citedMulti-Interest Network with Dynamic Routing for Recommendation at Tmall

51 citations · 87 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation

Niange Yu, Ye Tian, Biaolong Chen +5

Multi-subject video generation faces two key challenges: uncontrollable fidelity strength and potential semantic drift. We address these by analyzing the internal mechanisms of Dif…

cs.CV2026

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Shiyi Zhang, Mushui Liu, Yunze Tong +8

On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory signals, has achieved significant success in Large Language Mode…

cs.CV2026

RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation

Yuhan Li, Xianfeng Tan, Fangao Zeng +6

Standard clothing asset generation---restoring forward-facing flat-lay garment images from diverse real-world contexts---holds immense commercial value yet demands both macroscopic…

cs.CV2026

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

Yunze Tong, Mushui Liu, Canyu Zhao +9

With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align t…

cs.CV2026

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

Yuhan Li, Fangao Zeng, Sicong Kang +5

Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequenti…

cs.CV2026

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding

Peng Zhang, Guanghao Zhang, Wanggui He +10

Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video…