10 papers
TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation
Hongyu Zhang, Yufan Deng, Zilin Pan +6
Generating high-quality videos from complex temporal descriptions that contain multiple sequential actions is a key unsolved problem. Existing methods are constrained by an inheren…
Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models
Yuanyang Yin, Yufan Deng, Shenghai Yuan +3
The task of Image-to-Video (I2V) generation aims to synthesize a video from a reference image and a text prompt. This requires diffusion models to reconcile high-frequency visual c…
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
Hongyu Zhang, Yufan Deng, Shenghai Yuan +5
In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main c…
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
Yufan Deng, Yuanyang Yin, Xun Guo +8
We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual p…
MTPNet: Multi-Grained Target Perception for Unified Activity Cliff Prediction
Zishan Shu, Yufan Deng, Hongyu Zhang +2
Activity cliff prediction is a critical task in drug discovery and material design. Existing computational methods are limited to handling single binding targets, which restricts t…
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
Shenghai Yuan, Xianyi He, Yufan Deng +5
Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the…