From the 1 of 6 linked papers with an AI index.
6 papers
FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring
Jiatong Li, Leo Liang, Linghe Kong +1
The paper introduces FreqForcing, a training‑free method that uses spectral self‑anchoring to counter low‑frequency energy drift and improve visual stability in autoregressive long…
DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models
Xinrui Shi, Kai Liu, Ziqing Zhang +3
Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations…
InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution
Ziqing Zhang, Kai Liu, Zheng Chen +5
Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long s…
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
Junxian Li, Kai Liu, Leyang Chen +7
Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting compute…
Fose: Fusion of One-Step Diffusion and End-to-End Network for Pansharpening
Kai Liu, Zeli Lin, Weibo Wang +2
Pansharpening is a significant image fusion task that fuses low-resolution multispectral images (LRMSI) and high-resolution panchromatic images (PAN) to obtain high-resolution mult…
UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
Kai Liu, Leyang Chen, Wenbo Li +5
Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. However, evaluations of unified multimodal models (UMMs) rem…