7 papers
TASR: Timestep-Aware Diffusion Model for Image Super-Resolution
Qinwei Lin, Xiaopeng Sun, Yu Gao +4
Diffusion models have recently achieved outstanding results in the field of image super-resolution. These methods typically inject low-resolution (LR) images via ControlNet.In this…
Advancing Visual Large Language Model for Multi-granular Versatile Perception
Wentao Xiang, Haoxian Tan, Cong Wei +3
Perception is a fundamental task in the field of computer vision, encompassing a diverse set of subtasks that can be systematically categorized into four distinct groups based on t…
Optimizing Singular Spectrum for Large Language Model Compression
Dengjie Li, Tiancheng Shen, Yao Zhou +7
Large language models (LLMs) have demonstrated remarkable capabilities, yet prohibitive parameter complexity often hinders their deployment. Existing singular value decomposition (…
HiMix: Reducing Computational Complexity in Large Vision-Language Models
Xuange Zhang, Dengjie Li, Bo Liu +7
Benefiting from recent advancements in large language models and modality alignment techniques, existing Large Vision-Language Models(LVLMs) have achieved prominent performance acr…
Manga Generation via Layout-controllable Diffusion
Siyu Chen, Dengjie Li, Zenghao Bao +4
Generating comics through text is widely studied. However, there are few studies on generating multi-panel Manga (Japanese comics) solely based on plain text. Japanese manga contai…
LinVT: Empower Your Image-level Large Language Model to Understand Videos
Lishuai Gao, Yujie Zhong, Yingsen Zeng +3
Large Language Models (LLMs) have been widely used in various tasks, motivating us to develop an LLM-based assistant for videos. Instead of training from scratch, we propose a modu…