Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution
Wenxue Li, Jingjing Ren, Peng Zhang +4
High-resolution video generation faces a coupled bottleneck of optimization instability and prohibitive computational costs. The massive expansion of the token sequence not only bi…
cs.CV2025
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
Juntian Zhang, Chuanqi cheng, Yuhan Liu +3
Vision-language models (VLMs) achieve remarkable success in single-image tasks. However, real-world scenarios often involve intricate multi-image inputs, leading to a notable perfo…