3 papers
cs.CV2026
TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications
Feibo Jiang, Siwei Tu, Li Dong +5
Visual-Language Models (VLMs), with their strong capabilities in image and text understanding, offer a solid foundation for intelligent communications. However, their effectiveness…
cs.CV2025
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
Xiaolong Li, Youping Gu, Xi Lin +2
Attention mechanisms are the core of foundation models, but their quadratic complexity remains a critical bottleneck for scaling. This challenge has driven the development of effic…
cs.CV2025
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
Youping Gu, Xiaolong Li, Yuhao Hu +2
Diffusion Transformers currently lead the field in high-quality video generation, but their slow iterative denoising process and prohibitive quadratic attention costs for long sequ…