5 papers
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
Peiliang Cai, Evelyn Zhang, Jiacheng Liu +8
Recent advances in autoregressive video diffusion have enabled sequential and streaming video generation. However, long-horizon generation requires increasingly large KV caches, ma…
D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning
Evelyn Zhang, Fufu Yu, Aoqi Wu +5
Processing long visual token sequences poses a significant computational burden on Multimodal Large Language Models (MLLMs). While token pruning offers a path to acceleration, we f…
Estimation of Kinematic Motion from Dashcam Footage
Evelyn Zhang, Alex Richardson, Jonathan Sprinkle
The goal of this paper is to explore the accuracy of dashcam footage to predict the actual kinematic motion of a car-like vehicle. Our approach uses ground truth information from t…
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
Chang Zou, Evelyn Zhang, Shikang Zheng +6
Diffusion Transformers (DiT) have become the dominant methods in image and video generation yet still suffer substantial computational costs. As an effective approach for DiT accel…
Token Pruning for Caching Better: 9 Times Acceleration on Stable Diffusion for Free
Evelyn Zhang, Bang Xiao, Jiayi Tang +5
Stable Diffusion has achieved remarkable success in the field of text-to-image generation, with its powerful generative capabilities and diverse generation results making a lasting…