3 papers
cs.CV2026
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
Sihan Cao, Jianwei Zhang, Pengcheng Zheng +7
Large Vision-Language Models (LVLMs) incur substantial inference costs due to the processing of a vast number of visual tokens. Existing methods typically struggle to model progres…
cs.CV2025
Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor
Xiaoliu Guan, Lielin Jiang, Hanqi Chen +6
Diffusion Transformers (DiTs) have demonstrated remarkable performance in visual generation tasks. However, their low inference speed limits their deployment in low-resource applic…
cs.CV2025
Enabling Versatile Controls for Video Diffusion Models
Xu Zhang, Hao Zhou, Haoming Qin +5
Despite substantial progress in text-to-video generation, achieving precise and flexible control over fine-grained spatiotemporal attributes remains a significant unresolved challe…