works on

From the 1 of 21 linked papers with an AI index.

collaborators

21 papers

cs.CV2026

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Zizhong Ding, Junxian Li, Kai Liu +4

Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive e…

cs.CV2026

FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

Jiatong Li, Leo Liang, Linghe Kong +1

The paper introduces FreqForcing, a training‑free method that uses spectral self‑anchoring to counter low‑frequency energy drift and improve visual stability in autoregressive long…

cs.CV2026

InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution

Ziqing Zhang, Kai Liu, Zheng Chen +5

Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long s…

cs.CV2026

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks

Junxian Li, Kai Liu, Leyang Chen +7

Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting compute…

cs.CV2026

DVD-Quant: Data-free Video Diffusion Transformers Quantization

Zhiteng Li, Hanxuan Li, Junyi Wu +6

Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture for video generation, yet their computational and memory demands hinder practical deployment. While…

cs.CV2026

LSGQuant: Layer-Sensitivity Guided Quantization for One-Step Diffusion Real-World Video Super-Resolution

Tianxing Wu, Zheng Chen, Cirou Xu +5

One-Step Diffusion Models have demonstrated promising capability and fast inference in video super-resolution (VSR) for real-world. Nevertheless, the substantial model size and hig…