From the 1 of 23 linked papers with an AI index.
12 papers · 1 filter
LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
Enhuai Liu, Yunke Wang, Yutong Wang +2
Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and dura…
PARE: Pruning and Adaptive Routing for Efficient Video Generation
Yutong Wang, Yunke Wang, Tianfan Xue +4
Video Diffusion Transformers (DiTs) generate high-quality videos but demand substantial compute due to wide blocks, deep architectures, and iterative sampling. Recent methods reduc…
BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation
Yutong Wang, Yunke Wang, Xinyuan Chen +1
Automatic movie trailer generation must select shots from a full-length film and synchronize them with background music. Existing methods either relegate music alignment to post-pr…
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
Jingxuan He, Xiyu Wang, Yunke Wang +2
Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural language instructions, which impl…
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
Chenyu Hui, Xiaodi Huang, Siyu Xu +5
Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often s…
Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing
Jingxuan He, Xiyu Wang, Mengyu Zheng +3
Instruction-based image editing (IIE) aims to modify images according to textual instructions while preserving irrelevant content. Despite recent advances in diffusion transformers…