activity
20242026
collaborators

10 papers

cs.CV2026

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

Cong Han, Xiaohan Lan, Haibo Qiu +1

Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pivotal research frontier. The…

cs.CV2025

TASR: Timestep-Aware Diffusion Model for Image Super-Resolution

Qinwei Lin, Xiaopeng Sun, Yu Gao +4

Diffusion models have recently achieved outstanding results in the field of image super-resolution. These methods typically inject low-resolution (LR) images via ControlNet.In this…

cs.CV2025

Advancing Visual Large Language Model for Multi-granular Versatile Perception

Wentao Xiang, Haoxian Tan, Cong Wei +3

Perception is a fundamental task in the field of computer vision, encompassing a diverse set of subtasks that can be systematically categorized into four distinct groups based on t…

eess.IV2025

360-Degree Video Super Resolution and Quality Enhancement Challenge: Methods and Results

Ahmed Telili, Wassim Hamidouche, Ibrahim Farhat +12

Omnidirectional (360-degree) video is rapidly gaining popularity due to advancements in immersive technologies like virtual reality (VR) and extended reality (XR). However, real-ti…

cs.CL2025

Optimizing Singular Spectrum for Large Language Model Compression

Dengjie Li, Tiancheng Shen, Yao Zhou +7

Large language models (LLMs) have demonstrated remarkable capabilities, yet prohibitive parameter complexity often hinders their deployment. Existing singular value decomposition (…

cs.CV2025

HiMix: Reducing Computational Complexity in Large Vision-Language Models

Xuange Zhang, Dengjie Li, Bo Liu +7

Benefiting from recent advancements in large language models and modality alignment techniques, existing Large Vision-Language Models(LVLMs) have achieved prominent performance acr…