10 papers
AIR: Adaptive Interleaved Reasoning with Code in MLLMs
Cong Han, Xiaohan Lan, Haibo Qiu +1
Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pivotal research frontier. The…
TASR: Timestep-Aware Diffusion Model for Image Super-Resolution
Qinwei Lin, Xiaopeng Sun, Yu Gao +4
Diffusion models have recently achieved outstanding results in the field of image super-resolution. These methods typically inject low-resolution (LR) images via ControlNet.In this…
Advancing Visual Large Language Model for Multi-granular Versatile Perception
Wentao Xiang, Haoxian Tan, Cong Wei +3
Perception is a fundamental task in the field of computer vision, encompassing a diverse set of subtasks that can be systematically categorized into four distinct groups based on t…
360-Degree Video Super Resolution and Quality Enhancement Challenge: Methods and Results
Ahmed Telili, Wassim Hamidouche, Ibrahim Farhat +12
Omnidirectional (360-degree) video is rapidly gaining popularity due to advancements in immersive technologies like virtual reality (VR) and extended reality (XR). However, real-ti…
Optimizing Singular Spectrum for Large Language Model Compression
Dengjie Li, Tiancheng Shen, Yao Zhou +7
Large language models (LLMs) have demonstrated remarkable capabilities, yet prohibitive parameter complexity often hinders their deployment. Existing singular value decomposition (…
HiMix: Reducing Computational Complexity in Large Vision-Language Models
Xuange Zhang, Dengjie Li, Bo Liu +7
Benefiting from recent advancements in large language models and modality alignment techniques, existing Large Vision-Language Models(LVLMs) have achieved prominent performance acr…