From the 1 of 21 linked papers with an AI index.
21 papers
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs
Zizhong Ding, Junxian Li, Kai Liu +4
Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive e…
FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring
Jiatong Li, Leo Liang, Linghe Kong +1
The paper introduces FreqForcing, a training‑free method that uses spectral self‑anchoring to counter low‑frequency energy drift and improve visual stability in autoregressive long…
InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution
Ziqing Zhang, Kai Liu, Zheng Chen +5
Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long s…
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
Junxian Li, Kai Liu, Leyang Chen +7
Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting compute…
DVD-Quant: Data-free Video Diffusion Transformers Quantization
Zhiteng Li, Hanxuan Li, Junyi Wu +6
Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture for video generation, yet their computational and memory demands hinder practical deployment. While…
LSGQuant: Layer-Sensitivity Guided Quantization for One-Step Diffusion Real-World Video Super-Resolution
Tianxing Wu, Zheng Chen, Cirou Xu +5
One-Step Diffusion Models have demonstrated promising capability and fast inference in video super-resolution (VSR) for real-world. Nevertheless, the substantial model size and hig…