collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion

Kunyang Li, Mubarak Shah, Yuzhang Shang

Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attention leads to quadratic comput…

cs.CV2026

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls

Abdul Mohaimen Al Radi, Kunyang Li, Yuzhang Shang +2

Foundation video models produce visually impressive results, but their use in embodied AI remains limited because they are primarily trained on natural language rather than low-lev…

cs.CV2026

PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache

Kunyang Li, Mubarak Shah, Yuzhang Shang

A unified autoregressive model is a Transformer-based framework that addresses diverse multimodal tasks (e.g., text, image, video) as a single sequence modeling problem under a sha…

cs.CV2025

GVD: Guiding Video Diffusion Model for Scalable Video Distillation

Kunyang Li, Jeffrey A Chan Santiago, Sarinda Dhanesh Samarasinghe +2

To address the larger computation and storage requirements associated with large video datasets, video dataset distillation aims to capture spatial and temporal information in a si…

cs.CV2025

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana +66

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cul…